Tool
AI Cost Calculator
Across four levers: context, cache, calls and attempts. Plus model choice, billing path and human rework. The result is the only number that matters, the cost per completed task.
The article behind it: Cutting AI Costs Without Cutting QualityConfigure a run
Where do we start?
3,700 tokens per conversation, documented by Anthropic.
Billing and path (default: none)
Result
Cost per completed task: €2.50. Monthly total: €25,049.48.
- Input (machine)0%€0.003212
- Output (machine)0%€0.001736
- Rework (human)100%€2.50
Claude Haiku 4.5 ranks 13 of 27. Cheapest model: Ministral 3 3B at €2.50 per task, a gap of 0.18 percent.
Every model in this scenario
Calculated with 3,700 input and 400 output tokens, 1 calls, 1.0 attempts and 2 minutes of rework.
On machine price alone, the cheapest and the most expensive model are a factor of 139 apart. Per completed task the gap is 2.0 percent, because rework carries 99.8 percent of the cost. Model choice is the smaller lever here; the minutes are the bigger one.
Break-even: Ministral 3 3B only loses its lead over Claude Fable 5 once it needs 139.0 attempts where Claude Fable 5 gets by with 1.0. Check that number against your own failure rates before deciding on list price.
| Rank | Model | Machine | Human | Per task | Per month | Delta |
|---|---|---|---|---|---|---|
| 1 | Ministral 3 3B | €0.000356 | €2.50 | €2.50 | €25,003.56 | cheapest |
| 2 | DeepSeek V4 Flash | €0.000547 | €2.50 | €2.50 | €25,005.47 | +0.008% |
| 3 | Mistral Small 4 | €0.000690 | €2.50 | €2.50 | €25,006.90 | +0.013% |
| 4 | gpt-5.6-luna | €0.001059 | €2.50 | €2.50 | €25,010.59 | +0.028% |
| 5 | gpt-5.4-nano | €0.001076 | €2.50 | €2.50 | €25,010.76 | +0.029% |
| 6 | Gemini 3.1 Flash-Lite | €0.001324 | €2.50 | €2.50 | €25,013.24 | +0.039% |
| 7 | Qwen3.7-Plus (Together) | €0.001472 | €2.50 | €2.50 | €25,014.72 | +0.045% |
| 8 | DeepSeek V4 Pro | €0.001699 | €2.50 | €2.50 | €25,016.99 | +0.054% |
| 9 | Gemini 3.5 Flash-Lite | €0.001832 | €2.50 | €2.50 | €25,018.32 | +0.059% |
| 10 | Mistral Large 3 | €0.002127 | €2.50 | €2.50 | €25,021.27 | +0.071% |
| 11 | Llama 3.3 70B (Together) | €0.003701 | €2.50 | €2.50 | €25,037.01 | +0.13% |
| 12 | gpt-5.4-mini | €0.003971 | €2.50 | €2.50 | €25,039.71 | +0.14% |
| 13 | Claude Haiku 4.5· selected | €0.004948 | €2.50 | €2.50 | €25,049.48 | +0.18% |
| 14 | DeepSeek V4 Pro (Together) | €0.006797 | €2.50 | €2.51 | €25,067.97 | +0.26% |
| 15 | Gemini 3.6 Flash | €0.007422 | €2.50 | €2.51 | €25,074.22 | +0.28% |
| 16 | Mistral Medium 3.5 | €0.007422 | €2.50 | €2.51 | €25,074.22 | +0.28% |
| 17 | Gemini 3.5 Flash | €0.007943 | €2.50 | €2.51 | €25,079.43 | +0.30% |
| 18 | Claude Sonnet 5 (Einführungspreis bis 31.08.2026) | €0.009896 | €2.50 | €2.51 | €25,098.96 | +0.38% |
| 19 | gpt-5.6-terra | €0.0106 | €2.50 | €2.51 | €25,105.90 | +0.41% |
| 20 | Gemini 3.1 Pro Preview | €0.0106 | €2.50 | €2.51 | €25,105.90 | +0.41% |
| 21 | gpt-5.4 | €0.0132 | €2.50 | €2.51 | €25,132.38 | +0.52% |
| 22 | Claude Sonnet 5 (ab 01.09.2026) | €0.0148 | €2.50 | €2.51 | €25,148.44 | +0.58% |
| 23 | Claude Sonnet 4.6 | €0.0148 | €2.50 | €2.51 | €25,148.44 | +0.58% |
| 24 | Claude Opus 5 | €0.0247 | €2.50 | €2.52 | €25,247.40 | +0.98% |
| 25 | Claude Opus 4.8 | €0.0247 | €2.50 | €2.52 | €25,247.40 | +0.98% |
| 26 | gpt-5.6-sol | €0.0265 | €2.50 | €2.53 | €25,264.76 | +1.0% |
| 27 | Claude Fable 5 | €0.0495 | €2.50 | €2.55 | €25,494.79 | +2.0% |
Comparison of runs
Save runs to compare settings side by side: the same task with and without cache, say, or with twice the rework.
Formula, assumptions and sources
Cost per task = machine × modifiers + rework. The machine has four line items: fresh input (× calls × attempts), one cache write, cache reads (× every further call) and output (× calls × attempts). Rework is minutes/60 × €/h.
The cache is calculated the way it is billed: written once per task, then read. That is why it costs more than it saves on a single call. The calculator uses each model’s published cache prices rather than a flat factor. Anthropic charges 1.25 times to write and one tenth to read, DeepSeek reads for about one percent, Mistral publishes no cache price and therefore gets no discount. Google’s hourly storage fee depends on runtime and is NOT included.
Modifiers: batch × 0.5 (only at vendors that offer it), US residency × 1.1, router × 1.055. Attempts multiply calls, not the cache write and not the rework, because what gets reviewed is the result, not every try.
Prices are vendor list prices, verified 31 July 2026 against the pricing pages linked below, without discounts. Conversion uses 1.152 USD/EUR (reference rate 31 July 2026); that rate is an assumption, not a daily quote. Among the presets, only the 3,700 tokens for a support conversation is empirically documented (source: Anthropic documentation). All other presets are plausible example values, not measurements.
Cross-check: Anthropic’s documentation puts 10,000 support conversations of 3,700 tokens each on Claude Haiku 4.5 at about 37 US dollars. That is exactly what this calculator returns with output, cache and rework set to zero.