# ModuleDesk AI Draft Quality — Baseline Evaluation

**Date:** 2026-07-06 08:17 UTC
**Schema:** `tenant_internal` | **N:** 13 tickets evaluated

## Cost Assumptions

| Model | In ($/1M) | Out ($/1M) |
|-------|-----------|------------|
| gpt-4o | $2.50 | $10.00 |
| gpt-4o-mini | $0.15 | $0.60 |
| text-embedding-3-small | $0.02 | — |

EUR/USD rate: 1.09 (1 EUR = $1.09)

## Aggregate Results

| Metric | Value |
|--------|-------|
| Tickets evaluated | 13 |
| Median semantic similarity | 0.700 |
| Mean semantic similarity | 0.647 |
| % Usable — old judge (vs final reply) | 0% |
| **% Usable — honest judge (inbound-only)** | **46%** |
| Avg tone match | 2.8/5 |
| **Avg cost/draft** | **€0.00537** |
| Total run cost | €0.06984 |
| Projected cost (50 tickets) | €0.2686 |
| Projected cost (200 tickets) | €1.0745 |

## Call-Level Token Breakdown

| Call | Count | Prompt tok | Completion tok | Cost (€) |
|------|-------|-----------|----------------|----------|
| chat:gpt-4o | 12 | 22470 | 1450 | €0.06502 |
| chat:gpt-4o-mini | 13 | 9691 | 1463 | €0.00214 |
| embed:text-embedding-3-small | 26 | 4436 | 0 | €0.00008 |
| honest_judge:gpt-4o-mini | 13 | 7222 | 556 | €0.00130 |
| judge:gpt-4o-mini | 13 | 6763 | 647 | €0.00129 |

## Per-Ticket Results

| Ticket | Product | Lang | Similarity | Usable | Tone | Confidence | Draft tok | Cost |
|--------|---------|------|-----------|--------|------|------------|-----------|------|
| 6283 | Smart CSV Export | es | 0.720 | NO | 3/5 | Low (0) | 1194 | €0.00370 |
| 6281 | Conversion Pixel Tracking +  | es | 0.570 | NO | 3/5 | High (94) | 2542 | €0.00694 |
| 6277 | Custom Audiences | en | 0.739 | NO | 2/5 | Medium (77) | 2038 | €0.00572 |
| 6352 | Product Videos - Youtube, Vi | fr | 0.319 | NO | 2/5 | Low (5) | 0 | €0.00034 |
| 6262 | Google Adwords Conversion Tr | en | 0.555 | NO | 2/5 | Low (0) | 1506 | €0.00457 |
| 6353 | Pixel Plus for Facebook: Eve | fr | 0.763 | NO | 4/5 | Medium (72) | 3442 | €0.00958 |
| 6350 | Estimated Delivery Date V3 - | en | 0.700 | NO | 3/5 | Low (0) | 939 | €0.00324 |
| 6206 | Translation Suggestions | en | 0.524 | NO | 3/5 | Low (0) | 788 | €0.00257 |
| 6169 | WhatsApp Contact | en | 0.570 | NO | 2/5 | Medium (62) | 2177 | €0.00598 |
| 6175 | Products Alert | en | 0.778 | NO | 3/5 | High (90) | 2579 | €0.00742 |
| 6347 | Products Feed: Catalogue & S | es | 0.767 | NO | 3/5 | High (89) | 3250 | €0.00924 |
| 5520 | Dynamic Ads for Facebook: Ev | en | 0.795 | NO | 3/5 | High (94) | 2695 | €0.00807 |
| 4416 | Wire Transfer Remainder - Sm | en | 0.617 | NO | 4/5 | Low (0) | 770 | €0.00248 |

## Top Divergence Reasons

- **ticket 6283** (Smart CSV Export): The draft lacks specific information about the module's functionality and does not mention the 'Listas CSV' tab, which is crucial for the user's understanding.
- **ticket 6281** (Conversion Pixel Tracking + Custom Audiences): The draft does not address the specific issue of the plugin update and instead focuses on a different topic regarding email policies.
- **ticket 6277** (Custom Audiences): The draft does not acknowledge the seller's vacation status and has a more formal tone compared to the actual reply.
- **ticket 6352** (Product Videos - Youtube, Vimeo...): The AI draft does not address the specific issue of updating the module and lacks key information present in the actual reply.
- **ticket 6262** (Google Adwords Conversion Tracking - Smart Modules): The AI draft does not address the compatibility of the module or offer modifications, which are key points in the actual reply.
- **ticket 6353** (Pixel Plus for Facebook: Events + CAPI + Pixel Catalog): The AI draft lacks specific details about the module's behavior and troubleshooting steps that are present in the actual reply.
- **ticket 6350** (Estimated Delivery Date V3 - Smart Modules): The AI draft requests specific information while the actual reply suggests sharing access without asking for additional details, leading to a different resolution path.
- **ticket 6206** (Translation Suggestions): The draft does not address the specific functionality of the module regarding translations and lacks the assurance about Bing's translation capabilities.
- **ticket 6169** (WhatsApp Contact): The AI draft does not address the specific situation of the Prestashop update and module changes mentioned in the actual reply.
- **ticket 6175** (Products Alert): The AI draft does not mention the product box for missing combinations or stock, which is key information in the actual reply.
- **ticket 6347** (Products Feed: Catalogue & Shop for Facebook & Insta): The AI draft does not address the specific configuration issue with the module and lacks key information about the server performance and the need for an update.
- **ticket 5520** (Dynamic Ads for Facebook: Events + CAPI + Catalogue): The AI draft does not mention the specific modules included in the pack and lacks details about the Pixel Plus and its features, which are crucial for the customer's understanding.
- **ticket 4416** (Wire Transfer Remainder - Smart Modules): The draft does not specify that only bank transfers are currently supported, which is key information in the actual reply.

## Confidence Calibration (Phase 3+4)

| Band | Count | Mean sim | Old usable | Honest usable | Ask-for-info |
|------|-------|----------|------------|---------------|--------------|
| High | 4 | 0.727 | 0% | 1/4 | 1/4 |
| Medium | 3 | 0.690 | 0% | 1/3 | 0/3 |
| Low | 6 | 0.572 | 0% | 4/6 | 2/6 |

_Ask-for-info total: 3/13_
_Low-confidence → ask-for-info: 2 (Phase 4 target: Low drafts should ask, not invent)_

## Notes

- Semantic similarity is cosine distance between `text-embedding-3-small` embeddings of the draft
  and the actual sent reply (range 0–1, higher is better; ≥0.85 is strong, 0.70–0.85 is usable).
- Judge prompt asks gpt-4o-mini whether the seller could send the draft with **minor edits only**.
- Cost does NOT include the evaluation embeddings or judge calls in the "per-draft" figure
  (those are eval overhead). The per-draft cost covers classify + draft + RAG embedding.
- The classify+draft token counts come from the actual AIService calls; the `embed` rows reflect
  the query embedding computed during RAG context retrieval.
