Back to insights

Tool

AI Cost Calculator

Across four levers: context, cache, calls and attempts. Plus model choice, billing path and human rework. The result is the only number that matters, the cost per completed task.

The article behind it: Cutting AI Costs Without Cutting Quality

Configure a run

Where do we start?

3,700 tokens per conversation, documented by Anthropic.

1 · Context per call
2 · Cache hit rate
3 · Calls and attempts
4 · Model
5 · Human rework
Billing and path (default: none)

Result

Cost per completed task: €2.50. Monthly total: €25,049.48.

Cost per completed task
€2.50
Monthly total at 10,000 tasks: €25,049.48
  • Input (machine)0%€0.003212
  • Output (machine)0%€0.001736
  • Rework (human)100%€2.50
List prices in USD, converted at 1.152 USD/EUR (reference rate 31 July 2026).

Claude Haiku 4.5 ranks 13 of 27. Cheapest model: Ministral 3 3B at €2.50 per task, a gap of 0.18 percent.

Every model in this scenario

Calculated with 3,700 input and 400 output tokens, 1 calls, 1.0 attempts and 2 minutes of rework.

On machine price alone, the cheapest and the most expensive model are a factor of 139 apart. Per completed task the gap is 2.0 percent, because rework carries 99.8 percent of the cost. Model choice is the smaller lever here; the minutes are the bigger one.

Break-even: Ministral 3 3B only loses its lead over Claude Fable 5 once it needs 139.0 attempts where Claude Fable 5 gets by with 1.0. Check that number against your own failure rates before deciding on list price.

Every model, sorted by cost per completed task
RankModelMachineHumanPer taskPer monthDelta
1Ministral 3 3B€0.000356€2.50€2.50€25,003.56cheapest
2DeepSeek V4 Flash€0.000547€2.50€2.50€25,005.47+0.008%
3Mistral Small 4€0.000690€2.50€2.50€25,006.90+0.013%
4gpt-5.6-luna€0.001059€2.50€2.50€25,010.59+0.028%
5gpt-5.4-nano€0.001076€2.50€2.50€25,010.76+0.029%
6Gemini 3.1 Flash-Lite€0.001324€2.50€2.50€25,013.24+0.039%
7Qwen3.7-Plus (Together)€0.001472€2.50€2.50€25,014.72+0.045%
8DeepSeek V4 Pro€0.001699€2.50€2.50€25,016.99+0.054%
9Gemini 3.5 Flash-Lite€0.001832€2.50€2.50€25,018.32+0.059%
10Mistral Large 3€0.002127€2.50€2.50€25,021.27+0.071%
11Llama 3.3 70B (Together)€0.003701€2.50€2.50€25,037.01+0.13%
12gpt-5.4-mini€0.003971€2.50€2.50€25,039.71+0.14%
13Claude Haiku 4.5· selected€0.004948€2.50€2.50€25,049.48+0.18%
14DeepSeek V4 Pro (Together)€0.006797€2.50€2.51€25,067.97+0.26%
15Gemini 3.6 Flash€0.007422€2.50€2.51€25,074.22+0.28%
16Mistral Medium 3.5€0.007422€2.50€2.51€25,074.22+0.28%
17Gemini 3.5 Flash€0.007943€2.50€2.51€25,079.43+0.30%
18Claude Sonnet 5 (Einführungspreis bis 31.08.2026)€0.009896€2.50€2.51€25,098.96+0.38%
19gpt-5.6-terra€0.0106€2.50€2.51€25,105.90+0.41%
20Gemini 3.1 Pro Preview€0.0106€2.50€2.51€25,105.90+0.41%
21gpt-5.4€0.0132€2.50€2.51€25,132.38+0.52%
22Claude Sonnet 5 (ab 01.09.2026)€0.0148€2.50€2.51€25,148.44+0.58%
23Claude Sonnet 4.6€0.0148€2.50€2.51€25,148.44+0.58%
24Claude Opus 5€0.0247€2.50€2.52€25,247.40+0.98%
25Claude Opus 4.8€0.0247€2.50€2.52€25,247.40+0.98%
26gpt-5.6-sol€0.0265€2.50€2.53€25,264.76+1.0%
27Claude Fable 5€0.0495€2.50€2.55€25,494.79+2.0%

Comparison of runs

Save runs to compare settings side by side: the same task with and without cache, say, or with twice the rework.

Formula, assumptions and sources

Cost per task = machine × modifiers + rework. The machine has four line items: fresh input (× calls × attempts), one cache write, cache reads (× every further call) and output (× calls × attempts). Rework is minutes/60 × €/h.

The cache is calculated the way it is billed: written once per task, then read. That is why it costs more than it saves on a single call. The calculator uses each model’s published cache prices rather than a flat factor. Anthropic charges 1.25 times to write and one tenth to read, DeepSeek reads for about one percent, Mistral publishes no cache price and therefore gets no discount. Google’s hourly storage fee depends on runtime and is NOT included.

Modifiers: batch × 0.5 (only at vendors that offer it), US residency × 1.1, router × 1.055. Attempts multiply calls, not the cache write and not the rework, because what gets reviewed is the result, not every try.

Prices are vendor list prices, verified 31 July 2026 against the pricing pages linked below, without discounts. Conversion uses 1.152 USD/EUR (reference rate 31 July 2026); that rate is an assumption, not a daily quote. Among the presets, only the 3,700 tokens for a support conversation is empirically documented (source: Anthropic documentation). All other presets are plausible example values, not measurements.

Cross-check: Anthropic’s documentation puts 10,000 support conversations of 3,700 tokens each on Claude Haiku 4.5 at about 37 US dollars. That is exactly what this calculator returns with output, cache and rework set to zero.