Tool
AI Cost Calculator
Requirement first, then the number: pick the case, set in the requirement profile what an error costs and who reviews it, and see the cost per solved task for 86 models, with a range and the comparison to the literature.
The article behind it: Cutting AI Costs Without Cutting QualityQuestion
What do you want to calculate?
Background and evidence for: Answering a support ticket
3,700 tokens per conversation from the Anthropic documentation; 88 percent of AI-using companies deploy AI in customer contact (Bitkom, 02/2026).
Typical failure modes: A wrong answer that binds the company. Then the volume boomerang: Commonwealth Bank reversed 45 redundancies because the voice bot increased call volumes. Klarna rolled back in 2025 citing quality, and was back at 853 full-time equivalents as of Q3 2025.
Tokens in the GPT-5 reference measure (English text); for German text the calculator applies a per-model factor of 1.46 to 2.6.
Requirements
What does the result have to withstand?
rework: The error binds human time: an escalation, a query, a correction. Assumed: €0.50 to €5.
Source: Market price per resolved support ticket 0.49 to 2.00 USD (Drag, MavenAGI 2026); Zellinger and Thomson (arXiv 2507.03834): from about 0.01 USD error cost the stronger model wins. As of 19 August 2026.
spot-checked: Humans check samples. Caution: in one study 35 to 45 percent of flawed AI drafts were sent unchanged. Assumed: 0.5 minutes per task.
Source: Biro et al., MedStar/Georgetown, npj Digital Medicine 24 April 2025 (automation bias). As of 10 September 2026.
paragraphs: Short context, every model copes. Up to 8,000 tokens.
Source: NoLiMa, ICML 2025 (arXiv 2502.05167, v3 of 9 July 2025). As of 10 September 2026.
A wrong answer first costs an escalation and a correction, so human time; that is zone 3. As soon as the answer becomes legally binding, the case belongs in zone 4: Moffatt v. Air Canada, 2024 BCCRT 149, ended at 812 CAD for one wrong sentence from the chatbot. Checking is by sample, thinking effort stays low, and the answer has to hold across a few paragraphs.
Result
Answering a support ticket, Typical conversation, 10,000 per month.
Requirements: zone 3 rework, spot-checked, brief, paragraphs, English.
The measurement covers only 2 current models; the rest count as not measured from zone 3.
Show basis
One task: 3,700 input and 400 output tokens, 1 call, attempts per model from the solve rate on the anchor τ²-bench, customer service with policy and tools (tau2-bench (sierra-research/tau2-bench), pass rate. The README of the benchmark repo carries no table and points to taubench.com instead; the figures come from the CodeSOTA aggregator table, retrieved 19 August 2026., as of 19 August 2026) with a retry correction of 1.2. Monthly values assume 10,000 such tasks. 3,700 tokens per conversation from the Anthropic documentation; 88 percent of AI-using companies deploy AI in customer contact (Bitkom, 02/2026).
gpt-5.2 · Workhorse class
Attempts matter here: the ranking uses 1 divided by the solve rate on the anchor τ²-bench, customer service with policy and tools.
For this case the attempts come from the solve rate; the break-even over attempts does not apply.
Against zone 1 with machine-checkable review, your profile costs €10,400.00 per month.
Average model costs across the 42 models in the field sit at 1.4 times the cheapest eligible model. Review and error costs stay out of this comparison, they hit every model alike.
Cost per completed task: €1.06. Monthly total: €10,570.35.
Solve rate 73%, meaning 1.4 attempts on average.
Suitable models for this task
Class floor and solve rate decide, then price. Solve rates of the 2 rated models range from 59 to 73 percent, and cost per solved task varies by a factor of 1.
Read the assessment
A standard ticket runs about 3,700 input tokens and a few hundred output tokens; at this size the price gap between the compact class and the top class is at its widest. Whether the compact class holds shows in your attempts: at one attempt the price gap wins, and as attempts rise the arithmetic tips at the break-even point.
Own assessment, not a benchmark. The numbers above are price arithmetic, not a quality measurement.
The five cheapest in this scenario
Calculated on the basis shown in the report.
- 1Ministral 3 3BEdgebelow class floor
- Per solved task
- €0.50per task
- Per month, 10,000 tasks
- €5,003.52
- Solve rate
- —
- Gap to the recommendation
- -52.7%
- 2Ministral 3 8BEdgebelow class floor
- Per solved task
- €0.50per task
- Per month, 10,000 tasks
- €5,005.28
- Solve rate
- —
- Gap to the recommendation
- -52.6%
- 3DeepSeek V4 Flash (Together)hosted at Together AIWorkhorsenot measured
- Per solved task
- €0.50per task
- Per month, 10,000 tasks
- €5,005.41
- Solve rate
- —
- Gap to the recommendation
- -52.6%
- 4Qwen3.8 Flash (Together)hosted at Together AICompactnot measured
- Per solved task
- €0.50per task
- Per month, 10,000 tasks
- €5,006.38
- Solve rate
- —
- Gap to the recommendation
- -52.6%
- 5GLM-5.3-Flash (Together)hosted at Together AIWorkhorsenot measured
- Per solved task
- €0.50per task
- Per month, 10,000 tasks
- €5,006.48
- Solve rate
- —
- Gap to the recommendation
- -52.6%
- 41●gpt-5.2Workhorse
- Per solved task
- €1.06per task
- Per month, 10,000 tasks
- €10,570.35
- Solve rate
- 73 %
- Gap to the recommendation
- —
| Rank | Model | Per month, 10,000 tasks (€) | Gapto the recommendation |
|---|---|---|---|
| 1 | Ministral 3 3BEdgebelow class floor | €5,003.52 | -52.7% |
| 2 | Ministral 3 8BEdgebelow class floor | €5,005.28 | -52.6% |
| 3 | DeepSeek V4 Flash (Together)hosted at Together AIWorkhorsenot measured | €5,005.41 | -52.6% |
| 4 | Qwen3.8 Flash (Together)hosted at Together AICompactnot measured | €5,006.38 | -52.6% |
| 5 | GLM-5.3-Flash (Together)hosted at Together AIWorkhorsenot measured | €5,006.48 | -52.6% |
| Recommendation | |||
| 41 | ●gpt-5.2Workhorse | €10,570.35 | — |
Calculated for English text.
Rows without a solve rate on the anchor use the attempts you set.
What the literature recommends
Workhorse class as the default, compact class as the floor
The floor is the fast class of the strong vendors. Once the answer reaches the customer, the workhorse class is the default. τ²-bench spreads the models from 36 to 79 percent solve rate, and a wrong answer binds the company: Moffatt v. Air Canada ended at 812 CAD for one sentence. For the classes below Haiku and Flash no customer service measurement exists, in either direction. Caching the policies cuts the bill further than a step down in class: a cache read at Anthropic costs one tenth of the input price (pricing documentation, as of 10 September 2026).
- Anthropic, Choosing the right model 19 August 2026
- OpenAI, Model selection guide 19 August 2026
- tau2-bench leaderboard, pass rate 19 August 2026
- Moffatt v. Air Canada, 2024 BCCRT 149 14 February 2024
As of 19 August 2026
The calculator and the literature arrive at the same class here.
- Anthropic, pricing documentation (3,700 tokens per support conversation) 18 August 2026
- Bitkom, study report on artificial intelligence, 02/2026 19 August 2026
- Moffatt v. Air Canada, 2024 BCCRT 149, 14 February 2024 14 February 2024
- ABC News Australia on Commonwealth Bank, 21 August 2025 (no address) 21 August 2025
- Customer Experience Dive on Klarna, 20 November 2025 (no address) 20 November 2025
Fine-tune
Fine-tune: model, attempts, cache, calls
For this case the attempts per model come from the solve rate on the anchor; the slider has no effect.
The fee falls due when credit is topped up and does not sit on the individual call; paying by card it comes to at least 0.80 USD per top-up, by cryptocurrency it is 5 percent. Other routes, the Vercel Gateway among them, charge nothing for it.
All models
All values in EUR per 1M tokens · list prices, as of 10 September 2026
Cheapest, Frontier class
Mistral Large 3
€0.4291 / €1.29 per 1M tokens
Cheapest, Workhorse class
DeepSeek V4 Flash (Together)
€0.1202 / €0.2403 per 1M tokens
Cheapest, Compact class
Mistral Small 4
€0.1287 / €0.5149 per 1M tokens
As of
Prices 10 September 2026
Rate 1.1652 USD/EUR, ECB 9 September 2026
Capability (ECI) 9 September 2026
Knowledge
How the costs arise. Every card carries its sources.
What is a token?
The rule of thumb is one token per four characters of English, roughly 750 words per 1,000 tokens. German packs more densely: compounds and umlauts split into more pieces, so the same content costs a little more.
Every price on this page refers to 1M tokens. That sounds like a lot and fills up fast: a single 500 kB PDF already runs about 125,000 tokens, an eighth of it.
Sources: Anthropic, pricing documentation · OpenAI, pricing page
As of 10 September 2026
Input and output
Input covers system instructions, prior conversation and attached documents, and it is billed again on every single call. A typical 10 kB web page runs about 2,500 tokens.
Output includes reasoning steps, even when you never see them. Tasks with long output, drafts for example, shift the largest cost item from input to output.
Sources: Anthropic, pricing documentation · OpenAI, pricing page
As of 10 September 2026
The cache
The per-model prices sit in the model table: Anthropic charges 1.25 times the input price to write and one tenth to read, OpenAI charges the same 1.25 times to write since the gpt-5.6 series, Mistral has read for one tenth since August 2026, DeepSeek for about three percent.
This has a consequence that gets overlooked: on a single call, caching is MORE expensive than none, because it is written and never read. The benefit starts with the second call of the same task.
Every model and every effort level keeps its own cache. Switching mid-task forces a fresh write.
Below a minimum length no vendor stores anything. Anthropic names 512 tokens for the five series, 1,024 for most of the rest, 2,048 for Opus 4.7 and Haiku 3.5, and 4,096 for Opus 4.6, Opus 4.5 and Haiku 4.5; OpenAI 1,024 from the gpt-5.6 series on; Google 4,096 for Gemini 3.8, 3.7, 3.6 and 3.5 Flash plus 3.1 Pro and 2,048 for the 2.5 ones, while naming no figure at all for the Flash-Lite models. Fall below it and you get no error, just the full bill. This calculator sets the cache share to zero in that case and says so at the control.
Anthropic also offers a cache with a one-hour lifetime at twice the input price instead of 1.25 times. It pays off only when longer pauses sit between calls; in continuous work the five-minute cache is cheaper. This calculator uses the five-minute rate.
Where no minimum length is shown, the vendor publishes none: Mistral, DeepSeek and the hosters name none, and neither does OpenAI for the models before the gpt-5.6 series. The calculator grants the full cache discount there, because it knows nothing better. That is not a statement that no limit exists.
Sources: Anthropic, pricing documentation · Mistral, prompt caching documentation · DeepSeek, pricing page
As of 10 September 2026
Calls per task
A simple chat is one call. Once tools, search steps or intermediate checks join in, the number multiplies, and with it the input that is billed anew on every call.
This value stays easy to overlook: token prices sit on the pricing page, the number of calls sits nowhere and has to be measured.
Own assessment, no external source.
As of 10 September 2026
Attempts until it holds
This is where a cheap model proves whether it is actually cheaper. In research from March 2026, the cheaper-listed model produced the higher total cost in 32 percent of model pairs, mostly through highly variable thinking tokens and more working steps per task.
The break-even in the result works this out for your scenario: the number of attempts at which the cheapest model loses its lead to the priciest.
Where an anchor supplies a solve rate, the calculator derives the attempts from it, with a correction. Yang (arXiv 2605.08563, 8 May 2026) shows that assuming independent tries puts pass@3 17.4 points too high, 98.6 percent against 81.2 percent: a second attempt inherits the failure cause of the first. The calculator therefore applies a factor of 1.2 to the attempts. That factor is a stated assumption drawn from this measurement, not a measurement of our own.
Sources: Chen et al. 2026, arXiv 2603.23971 · Yang, arXiv 2605.08563, 8 May 2026
As of 10 September 2026
The second price tier
The higher rate then applies to ALL tokens of the call, including those below the threshold. That is how the vendors bill, and the thresholds often sit in a footnote below the pricing table.
A calculator that flatly uses the page-one price therefore underestimates long documents by up to half. Anthropic does not tier and offers the full window at the base price.
The "Trait" filter in the model table shows which 10 models are affected; in the result, affected rows carry the marker "long-context rate".
Sources: OpenAI, pricing page · Google, Gemini pricing page
As of 10 September 2026
The context window
The range is wide: Claude Haiku 4.5 takes up to 200,000 tokens, the current Opus and gpt-5.6 models around 1M. In the result, models whose window is too small for your input carry the marker "does not fit".
A dash in the model table means the vendor publishes no size for this model on its pricing page. No estimated number stands in for it.
Sources: Anthropic, pricing documentation · OpenAI, pricing page
As of 10 September 2026
Batch processing
Anthropic, OpenAI, Google and Mistral publish the batch discount as its own price list; the open-weights providers in our set list none. Where it is missing, this calculator drops the discount and says so at the number.
For recurring volume work, overnight document processing for example, that is 50 percent off; quality stays the same and only the wait is added.
OpenAI has run a third tier between the two since 2026: Flex bills at batch rates but answers immediately. In exchange a request can be turned away when no capacity is free; rejected requests cost nothing. For the OpenAI models that offer Flex, the batch control therefore shows the Flex price as well; at every other vendor it still stands for the queue.
Sources: Anthropic, pricing documentation · OpenAI, pricing page · Google, Gemini pricing page
As of 10 September 2026
Routers and data residency
Routers apply no markup to the tokens themselves; top-ups carry 5.5 percent by card. Requests are forwarded to whichever provider is selected, and that provider may use them for training or improvement. A fixed subprocessor list requires pinning the routing down.
Forced US processing costs 1.1 times the standard rate at Anthropic. The residency parameter currently knows no EU value; EU processing runs through the cloud platforms offering European regions.
Sources: OpenRouter, pricing page · Anthropic, pricing documentation
As of 10 September 2026
The five model classes
The mapping follows vendor naming: Anthropic tiers Opus, Sonnet, Haiku; OpenAI sol, terra, luna plus the pro tiers; Google Pro, Flash, Flash-Lite; Mistral Large, Medium, Small, plus the Ministral line explicitly for on-device use and Codestral and Devstral for code.
The class sorts, the capability column measures. The two can diverge: DeepSeek V4 Pro is its vendor’s frontier class at a compact-class price, and an edge-class model can top list-price rankings while its small context window and missing capability measurement rule it out for many business cases.
Only models bookable at an active endpoint with a published list price enter the set. Circulating headline prices without a bookable endpoint stay out, however tempting the number looks.
Sources: Anthropic, pricing documentation · OpenAI, pricing page · Google, Gemini pricing page · Mistral, API pricing page
As of 10 September 2026
Capability, difficulty, solve rate
The Epoch Capabilities Index (ECI) aggregates over 50 benchmarks into one capability value per model via item response theory, the same tooling that calibrates exams such as the GMAT. The scale is relative: GPT-5 sits at 150 by definition, Claude 3.5 Sonnet at 130.
The economic consequence is cost-of-pass: cost per solved task equals cost per attempt divided by the solve rate. A model just below the task difficulty gets expensive through attempts; far above it you pay for capability the task does not need.
The calculator uses this for business cases with an anchor benchmark, and only with measured solve rates: where no measurement exists, a dash stands instead of an estimated number, and from zone 3 upwards the row cannot win. Where no anchor holds, the report says so openly.
Sources: Epoch AI, Epoch Capabilities Index (CC-BY) · Ho et al. 2025, A Rosetta Stone for AI Benchmarks, arXiv 2512.00193 · Erol et al. 2025, Cost-of-Pass, arXiv 2504.13359
As of 9 September 2026
Prices and comparability
Verified 10 September 2026 directly against vendor pricing pages, without discounts. Prices change silently: between two checks of this calculator, one model temporarily sat five times too high in the list, with no notice anywhere.
Anthropic documents a tokeniser for Claude 4.7 and later that produces roughly 30 percent more tokens for the same text. A price comparison across vendors therefore only holds as an order of magnitude.
Two movements run against each other: the price for a given level of performance falls five to ten times a year, while the running cost of the strongest models rises three to eighteen times a year (Gundlach et al., 2026). Every cost figure needs a date, this one included.
Sources: Anthropic, pricing documentation · OpenAI, pricing page · Gundlach et al., The Price of Progress, arXiv 2511.23455 (v2, 23 March 2026)
As of 10 September 2026
What an error costs
Zellinger and Thomson put a number on the threshold: “reasoning models offer better accuracy-cost tradeoffs as soon as the economic cost of a mistake exceeds $0.01”. Model cascades lose their advantage from about $0.1. One precondition belongs with it: the threshold holds as long as the stronger model is also the more accurate one. For code review and research the measurements often run the other way; the threshold logic still holds there, the ranking behind it has to be measured.
In customer service the two figures sit two orders of magnitude apart. A ticket of 3,700 tokens costs fractions of a cent in every model class, while the market price of a resolved ticket runs 0.49 to 2.00 USD (Drag and MavenAGI 2026, secondary sources). Saving on the token price while leaving out escalation optimises the smaller of the two numbers.
Two documented cases show the upper end. In Moffatt v. Air Canada (2024 BCCRT 149, 14 February 2024) the tribunal awarded 812.02 CAD and stated: “It should be obvious to Air Canada that it is responsible for all the information on its website.” Klarna announced in February 2024 the work of a calculated 700 full-time agents and $40 million in profit improvement, and corrected course in May 2025: “As cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality.”
The requirements section is where you set this figure as a zone. The presets run from 0 euro through the cost of one more call up to 100 euro per task. The top zone carries no figure, because the error cannot be priced there, and it switches the result to model group and process.
Sources: Zellinger and Thomson, Economic Evaluation of LLMs, arXiv 2507.03834 · Drag and MavenAGI 2026, market price per resolved ticket (secondary source) · Moffatt v. Air Canada, 2024 BCCRT 149, 14 February 2024 · Klarna, announcement 27 February 2024 and correction 05/2025 (secondary source)
As of 10 September 2026
Checking costs time and catches part of it
A simulation study by MedStar and Georgetown (npj Digital Medicine, 24 April 2025) put 18 patient messages with AI drafts in front of 20 primary care physicians, four of them carrying planted errors. Between 35 and 45 percent of the flawed drafts went out unchanged, every flawed draft was missed by at least 13 of the 20 participants, on average 2.67 of 4 errors went unnoticed, and exactly one participant addressed all four. The authors name the cause as automation bias, “the tendency to over-rely on automation”.
Checking costs measurable time. At UC San Diego reading time per message rose by 21.8 percent (JAMA Network Open, 15 April 2024, p equals 0.008), while reply time fell by 5.9 percent without statistical significance. The calculator therefore applies review minutes times an hourly rate, with 60 euro per hour as a stated assumption.
Attention is a finite resource. Turan (arXiv 2606.08919, 8 June 2026) models reviewers who tire as the escalation load grows: “when the reviewer is modeled as endogenous (fatiguing as escalation load grows), realized safety becomes an inverted-U in the escalation rate: more human oversight can make a system less safe”. Reviewers also agree only moderately on what counts as risky, Fleiss kappa 0.52. The work is a single-author preprint without a human study; it stands here as a caution and does not enter the calculation.
OpenAI writes about GDPval (25 September 2025) that frontier models handle those tasks around 100 times faster and around 100 times cheaper than industry experts, and clarifies in the same passage that these figures cover inference time and API cost only and leave out human oversight, revision and integration. Those are the items the reviewability control brings into the calculation.
Sources: MedStar and Georgetown, npj Digital Medicine, 24 April 2025 · UC San Diego (Tai-Seale et al.), JAMA Network Open, 15 April 2024 · Turan 2026, Oversight Has a Capacity, arXiv 2606.08919 · OpenAI, GDPval, 25 September 2025
As of 10 September 2026
The price reversal
Chen et al. (arXiv 2603.23971, 25 March 2026, 8 models, 12 tasks) measure: “in 32% of model-pair comparisons, the model with a lower listed price actually incurs a higher total cost, with reversal magnitude reaching up to 28x”. On the same query one model burns up to 900 percent more thinking tokens than another, or ten times as many turns. Repeating the same query on the same model, thinking tokens vary by up to 9.7 times; the paper calls that an irreducible noise floor for any predictor.
OckBench (arXiv 2511.05722) measures the same effect at equal accuracy: more than a 25-fold token difference, about 1,600 against about 42,000 tokens. The paper puts it this way: “cheaper token cost does not always imply cheaper task cost; verbose smaller models can pay an Overthinking Tax”.
The effort level works at the same magnitude. Artificial Analysis measured GPT-5 at high effort using 23 times the tokens of minimal, 82M against 3.5M for the whole index, at 68 against 44 index points; the step from medium to high added one point. The generation counts too: Claude Sonnet 5 produces around 40 percent more output tokens per index task than Sonnet 4.6, needs around three times as many agent turns and costs around twice as much per task, at the same list price.
This version therefore calculates with a volume factor per model and effort level and shows the result as a range. Where no measured spread exists for the upper band, that stands at the figure.
Sources: Chen et al. 2026, arXiv 2603.23971 · OckBench, arXiv 2511.05722 · Artificial Analysis, GPT-5 benchmarks and analysis, 7 August 2025 · Artificial Analysis, Claude Sonnet 5 agentic cost, 30 June 2026
As of 10 September 2026
The tokeniser as a volume factor
Anthropic documents a new tokeniser for models from Claude 4.7 on: “This tokenizer produces approximately 30% more tokens for the same text.” At the same price per token that is a surcharge of roughly 30 percent on the volume side, without any price figure changing.
Across vendors the range is wider. On 10 June 2026 TextKit measured German text at 1.71 tokens per word for GPT-5, 2.18 for GPT-4, 2.64 for Claude Sonnet 4.6 and 3.48 for Claude Opus 4.8; English text sits at 1.17 to 1.88. Between the ends lies a factor of 2.0, and it hits input and output alike.
Measured across 24 EU languages (arXiv 2605.24718) the token count per word spreads by a factor of 2.5, from 1.2 in English to 3.1 in Greek. German averages 1.76 and ranges from 1.55 to 1.98 depending on the vendor, so 1.28 times through the choice of vendor alone.
The calculator applies the factor per model family to input and output and shows it as its own column in the model table. A comparison that treats tokens as a vendor-neutral unit favours the vendors with the coarser tokeniser.
Sources: Anthropic, pricing documentation · TextKit, tokens per word across vendors, 10 June 2026 · Tokeniser surcharge across 24 EU languages, arXiv 2605.24718
As of 10 September 2026
Window size and context fidelity
Two vendors report the same long-context value in their model cards, MRCR v2 with eight needles, and that is where the label falls apart. Gemini 3.1 Pro drops from 84.9 percent at 128,000 tokens to 26.3 percent at 1M (model card, 19 February 2026). Anthropic writes about the 1M variant: “on the 8-needle 1M variant of MRCR v2 … Opus 4.6 scores 76%, whereas Sonnet 4.5 scores just 18.5%”. Same label, fourfold difference.
NoLiMa (ICML 2025) measures the drop where the question does not literally overlap with the passage: at 32,000 tokens eleven of thirteen models fall below half their short-context performance. RULER confirms this across 17 models and reports an effective length below the advertised one: Llama-3.1-70B is listed at 128k and carries 64k.
Chroma shows that the drop hits simple tasks as well: “model performance varies significantly as input length changes, even on simple tasks”. Needle search tests less than real analysis work, which sits above it.
The length at which a measurement was taken is part of the claim. Google reports 97.0 percent for Gemini 3.7 Flash and 91.8 for 3.6 Flash, both at 128,000 tokens and as a cumulative average; it publishes no value for the full million. Until 10 September 2026 that 97 sat here as a 1M value, taken from a secondary source. It was wrong, and the difference is not a detail: the only measured 1M value of any Gemini model is 26.3 percent.
In the model table fidelity has its own column for 128k and 1M. Five models in the set carry a value, all from vendor model cards; all others carry a dash, because no figure exists. The check runs against the amount of text that actually occurs: up to 128,000 tokens against the 128k value, above it against the one for the full length. A measured value below 50 percent excludes. Where the value is missing, the field decides: as long as another model passes the threshold on evidence, a model without a value cannot win the case. If none passes it, the gap hits all of them equally, and the calculator says so at the number instead of blocking every row.
Sources: Google DeepMind, Gemini 3.1 Pro model card, 19 February 2026 · Google DeepMind, Gemini 3.7 Flash model card, 13 August 2026 · Anthropic, Claude Opus 4.6, 5 February 2026 · NoLiMa, ICML 2025, arXiv 2502.05167 · NVIDIA RULER, effective context length · Chroma, Context Rot, 14 July 2025
As of 10 September 2026
The same question twice
On 10 September 2025 Thinking Machines Lab traced the cause to the varying batch size on the server: with the load, the order in which the kernel reduces changes as well. Measured, 1,000 identical requests at temperature 0 produced eighty different answers. With batch-invariant kernels all 1,000 answers were identical, and runtime rose from 26 to 55 seconds, and to 42 with an improved attention kernel.
The variation does not stay in the wording. Atil et al. (arXiv 2408.04667, five models, eight tasks, ten runs) report: “We see accuracy variations up to 15% across naturally occurring runs with a gap of best possible performance to worst possible performance up to 70%.” Yuan et al. (arXiv 2506.09501) measure “up to 9% variation in accuracy and 9,000 tokens difference in response length” on a reasoning model and trace it to floating-point arithmetic at limited precision.
Reliability therefore comes from the architecture around the model. The documented base pattern: the model proposes, a fixed rule or a checking program decides, and only the decision counts. Structured outputs secure the form of the answer; about the content they say nothing, and the vendor notes that such outputs can still contain mistakes. Repeated sampling and voting schemes lower the spread and do not create determinism, because they draw several samples on purpose.
The BSI lists non-reproducibility as a risk category of its own, R7 (Generative AI Models, version 2.0, 17 January 2025): “The outputs of many generative AI models are not necessarily reproducible due to the use of random components.” Measure M20 requires, where the impact may be critical, a review of the output, cross-referencing with further sources and manual post-processing where needed. In the result, the top error-cost zone shows the documented patterns with their limits.
Sources: Thinking Machines Lab, Defeating Nondeterminism in LLM Inference, 10 September 2025 · Atil et al., Non-Determinism of Deterministic LLM Settings, arXiv 2408.04667 · Yuan et al., Numerical Sources of Nondeterminism in LLM Inference, arXiv 2506.09501 · BSI, Generative AI Models, version 2.0, 17 January 2025, risk R7 and measure M20
As of 10 September 2026
Felt and measured effect
METR (Becker, Rush, Barnes, Rein, arXiv 2507.09089, 12 July 2025) had experienced developers work on mature code, randomised with and without AI. Measured, they were 19 percent slower with AI, while they had estimated themselves faster. Between self-assessment and measurement lie around 39 percentage points.
The counter-figures belong with it. A rollout across tens of thousands of developers at Microsoft shows around 24 percent more merged pull requests (arXiv 2607.01418, 1 July 2026). And the Stanford AI Index 2026 records that METR later could not replicate its own finding, “primarily due to a growing reluctance among developers to work without AI, and that developers in late 2025 were likely sped up by AI relative to the original study period” (page 219).
One rule for this calculator follows: it is never calibrated with self-assessments. Every case study resting on a satisfaction survey can be optimistic by that order of magnitude, and vendor figures on time saved carry the same risk. The numbers here come from price lists, model cards and measurement series with a published setup.
Sources: Becker, Rush, Barnes, Rein (METR), arXiv 2507.09089, 12 July 2025 · Microsoft rollout across tens of thousands of developers, arXiv 2607.01418, 1 July 2026 · Stanford AI Index 2026, page 219 (note on the METR replication)
As of 10 September 2026
Which benchmark per business case
GPQA Diamond is built to be “Google-proof”, that is, built so searching does not help. A measure that neutralises search ability cannot assess it. On top of that it is saturated: on 15 August 2026 vals.ai counts “24 of the 133 models tested now score 90% or higher”, and the benchmark author, David Rein, says in an interview that it “stopped discriminating between good and great”. For research, BrowseComp and Deep Research Bench hold, the latter being the only one that reports cost per task.
SWE-bench Verified measures producing a patch when the fault location is known. A review has to find the fault first. CR-Bench (arXiv 2603.11078, 10 March 2026) sets out the difference and introduces its own measures, usefulness rate and signal-to-noise: “these agents often face a fundamental trade-off, they either prioritize precision and risk missing critical vulnerabilities, or prioritize recall at the cost of producing noisy and low-actionable feedback.”
For support tickets τ²-bench holds, because it tests the three properties of the case: a policy the system has to follow, tool use, and a simulated customer who follows up on incomplete answers. The spread on the public leaderboard is wide: Claude Opus 4.5 79 percent, GPT-5.2 73, Sonnet 4.5 63, GPT-4o 36. As a blocking criterion for source-bound output the HHEM hallucination leaderboard serves, read as a minimum threshold, because a low rate can also reflect low capability.
For text work there is a retrievable anchor with a price column, arena.ai creative writing (as of 12 August 2026, 1.18M votes). Confidence intervals of ±7 to ±18 points in the top group (±5 to ±28 across the whole field) make the leading ranks statistically indistinguishable. The anchor therefore works as a threshold: it answers whether a model belongs to the top group, and inside that group the price decides.
Sources: MindStudio, interview with GPQA author David Rein, 5 May 2026 · vals.ai, GPQA evaluation, 15 August 2026 · Deep Research Bench, as of 10 June 2026 · CR-Bench, arXiv 2603.11078, 10 March 2026 · τ²-bench, sierra-research/tau2-bench · CodeSOTA, τ²-bench leaderboard 2026 (secondary source) · Vectara HHEM-2.3, values reported via CodingFleet 2026 (secondary source) · arena.ai, creative writing, as of 12 August 2026
As of 10 September 2026
The five error zones
Error resilience in the requirement profile has five levels. Each level carries an amount per task and a source; zone 5 carries no figure, there the process decides.
- Zone 1: doesn't matter (€0)An error costs nothing; the result is discarded at once or overwritten anyway.Source: Anthropic, Optimizing for cost and intelligence: "high-volume work with checkable outputs"
- Zone 2: cheap to fix (about one call)The error is noticed and costs one more call, nothing else.Source: Erol et al., Cost-of-Pass (arXiv 2504.13359, v2 of 26 February 2026); Anthropic, Optimizing for cost and intelligence: "Run everything at low effort and re-run failures at the default (high); on the coding benchmark measured, the pass rate held at about half the cost"
- Zone 3: rework (€0.50 to €5)The error binds human time: an escalation, a query, a correction.Source: Market price per resolved support ticket 0.49 to 2.00 USD (Drag, MavenAGI 2026); Zellinger and Thomson (arXiv 2507.03834): from about 0.01 USD error cost the stronger model wins
- Zone 4: business impact (€50 to €500)The error costs customers, deadlines or legal positions.Source: Cost-of-Pass, human reference GPQA Diamond 58 USD per task (arXiv 2504.13359); Moffatt v. Air Canada 2024 BCCRT 149 (812 CAD); Klarna rollback 05/2025
- Zone 5: must not happen (no figure, blocked)The error cannot be priced or is existential. A rule set in advance decides.Source: BSI Generative AI Models v2.0 (17 January 2025) M20; AI Act Regulation (EU) 2024/1689 Art. 14, 15; Lemonade: "we never let AI perform deterministic actions"
Formula, assumptions and sources
1 · Formula
Cost per task = machine × modifiers. The machine has four line items: fresh input (× calls × attempts), one cache write, cache reads (× every further call) and output (× calls × attempts). The cache write counts once per task and is not multiplied by attempts, because it is billed only once as well. Human rework is deliberately not part of this calculator. It compares models.
Since version 3 the expected calculation per solved task sits above it: K = cost per attempt × attempts + (1 − solve rate) × error costs + review costs. The attempt is the machine calculation above, extended by the volume and tokeniser factors. Attempts come from the solve rate at the case anchor, times 1.2, because repeated tries inherit the same failure cause: Yang (arXiv 2605.08563) measures that assuming independent tries overstates the hit rate after three attempts by 17.4 points. The direction is evidenced, the size is not: the factor 1.2 is a stated assumption of this calculator. Error costs are the amount of the selected zone and fall due at the counter-probability to the solve rate. Review costs are review minutes times an hourly rate, assumed at 60 euro per hour. The machine calculation from v2 stays untouched: with error costs 0, review costs 0, normal effort, a neutral tokenizer factor and no anchor, the formula returns exactly the value of the earlier version.
2 · Cache
The cache is calculated the way it is billed: written once per task, then read. That is why it costs more than it saves on a single call. The calculator uses each model’s published cache prices. Anthropic charges 1.25 times to write and one tenth to read, OpenAI charges the same 1.25 times to write since the gpt-5.6 series, DeepSeek reads for about three percent, Mistral has read for one tenth of the input price since August 2026. Google’s hourly storage fee depends on runtime and is NOT included.
3 · Surcharges and discounts
Modifiers: batch × 0.5 (only at vendors that offer it), US residency × 1.1, router × 1.055. Attempts multiply calls. Running a task twice still writes the cache once.
4 · Second price tier
10 models switch to a more expensive tier above a certain context size: Gemini 2.5 Pro and Gemini 3.1 Pro above 200,000 tokens, gpt-6-astra, the gpt-5.6 series, gpt-5.5, gpt-5.5-pro, gpt-5.4 and gpt-5.4-pro above 272,000. The higher rate then applies to every token of the call, including those below the threshold, because that is how the vendors bill it. Affected rows carry the note "long-context rate" in the ranking. Anthropic does not tier and offers the full window at the base price.
5 · Capability and solve rate
The section "All models" shows the Epoch Capabilities Index per model (Epoch AI, data CC-BY, as of 9 September 2026), a capability value merged across 50-plus benchmarks via item response theory. The scale is relative: GPT-5 sits at 150 by definition, Claude 3.5 Sonnet at 130.
9 of the 12 business cases carry an anchor benchmark from anker.json, each with source, licence and date. Only anchors of type "rate" yield a solve rate per model; attempts follow from it under the cost-of-pass principle: cost per solved task equals cost per attempt divided by solve rate. Anchors of type "threshold" (Elo), "gate" (hallucination rate) and "fidelity" (MRCR, context fidelity across the full window) only filter and are never read as an attempt count; for "Working through a long document" the attempts therefore come from the case's manual value.
Nothing is estimated: what sits in anker.json is measured, and models without a value show a dash. From zone 3 upwards a row without a measurement cannot win. Rates are bounded to 5 through 99 percent. The benchmark approximates the business case; it does not measure the concrete task, and expert mode keeps calculating with manual values. Where a case carries no anchor, the assessment says what the ranking follows instead.
6 · Model set
The set covers 86 models, including deprecated and retired ones. The ranking in the result hides them by default, because nobody should adopt them fresh; anyone running an existing estate brings them back through the checkbox. The section "All models" does the same: the table shows the current ones by default, and the switch "show superseded and retired" above it brings in the rest. Where the context window is too small for the configured input, the row reads "does not fit". A cheap model is no use if the task will not fit. Every model carries one of five classes (frontier, workhorse, compact, edge, code), derived from the vendor’s own positioning; the class sorts, measurement sits in the capability column. Only models bookable at an active endpoint with a published list price enter the set; circulating headline prices without a bookable endpoint stay out.
7 · Evidence depth
Every price, cache rate, batch discount and context tier comes straight from the vendor pricing pages linked below. A context window is published for 83 of the 86 models; only Claude Opus 4, Claude Sonnet 4 and Claude Haiku 3.5 carry a dash, because the vendor publishes no size for them. The deprecated classification follows the vendor’s own model overview at Anthropic; at the other vendors, which publish no such marking, it is my own reading by model generation.
8 · Price status and conversion
Prices are vendor list prices, verified 10 September 2026 against the pricing pages linked below, without discounts. Conversion uses 1.1652 USD/EUR (ECB reference rate of 9 September 2026); that rate is an assumption, not a daily quote. Among the business cases, only the 3,700 tokens for a support conversation is empirically documented (source: Anthropic documentation). All other sizes are plausible example values, not measurements. The dates are deliberately separate: prices as of 10 September 2026, anchors and zone evidence as of 10 September 2026.
9 · Cross-check
Anthropic’s documentation puts 10,000 support conversations of 3,700 tokens each on Claude Haiku 4.5 at about 37 US dollars. That is exactly what this calculator returns in expert mode with output and cache set to zero.
Sources
- Anthropic: pricing documentation (prices, cache rates, context windows, cross-check)
- OpenAI: pricing page (prices, tiers from 272,000 tokens)
- Google: Gemini pricing page (prices, tiers from 200,000 tokens, storage fee)
- Mistral: API pricing (prices, cache discount)
- DeepSeek: pricing page
- Together AI: pricing page (open weights at a hoster)
- OpenRouter: pricing page (router fee)
- ECB: euro reference rate USD
- Epoch AI: Epoch Capabilities Index (capability values and solve rates, data CC-BY)
- Ho et al. 2025: A Rosetta Stone for AI Benchmarks (method behind the capability index, arXiv 2512.00193)
- Erol et al. 2026: Cost-of-Pass (cost per solved task, arXiv 2504.13359, v2 of 26 February 2026, ICLR 2026)
- Chen et al. 2026: The Price Reversal Phenomenon (price reversal between list price and total cost, arXiv 2603.23971, v2 of 28 May 2026)
Changes
Prices, anchors, capability values and the exchange rate age independently of one another, and none of these figures announces itself when it goes wrong. This is what moved between two dates and where the new figure comes from. Latest: 10 September 2026.
FixedPrices
Since 21 August 2026 gpt-5.6-sol has been on a promotional price of 4 instead of 5 USD per million input tokens and 20 instead of 30 on output; the long tier sits at 8 and 30 instead of 10 and 45. The set carried the old rate and priced the model up to 50 percent too high. OpenAI calls the price valid "at least through 21 November 2026" and publishes no follow-on rate.
Sources: OpenAI, Preisseite
AddedModels
Four new frontier models in the set: Claude Fable 5.1 and Claude Mythos 5.1 (1 September), gpt-6-astra (3 September) and Gemini 3.8 Flash (2 September). Fable 5.1 brings its own cache read rate of 0.25 USD, a quarter of the usual tenth.
Sources: Anthropic, Preisdokumentation · OpenAI, Preisseite · Google, Gemini-Preisseite
AddedModels
Seven new endpoints at the hoster Together: DeepSeek V4 Flash, Kimi K3, GLM-5.3, GLM-5.3-Flash, GLM-5.2, Qwen3.8-2.4T-A95B and Qwen3.8 Flash, all with a one-million-token window. DeepSeek V4 Flash costs 0.14 instead of 0.44 USD there, making the hoster cheaper than the model provider.
Sources: Together AI, Preisseite
ChangedPrices
Together has moved DeepSeek V4 Pro to 1.32 and 3.96 USD, previously 1.74 and 3.48, and reads from cache at 0.13 instead of 0.20. The context window of Qwen3.7-Max now stands at one million instead of 262,144 tokens. The batch discount at Together applies to Llama 3.3 70B only; every other model there bills at the full rate.
Sources: Together AI, Preisseite · Together AI, Batch-Dokumentation
AddedMethod
Minimum cache lengths are now part of the calculation. Below a certain length no vendor stores anything, and does so without an error: Claude Haiku 4.5 and the Gemini 3 Flash models plus Gemini 3.1 Pro require 4,096 tokens, OpenAI 1,024 from the gpt-5.6 series on, the Claude five series 512. On short cases such as classification or email the part to be cached fell below that, and the calculator showed a discount that does not exist.
Sources: Anthropic, Prompt Caching · OpenAI, Prompt Caching · Google, Context Caching
FixedAnchors
The context fidelity of Gemini 3.7 Flash stood at 97 percent as a value for the full window length. Google’s model card reports it at 128,000 tokens, and as a cumulative average; for one million tokens Google publishes nothing. The old entry came from a secondary source. As a result no current model carries an evidenced fidelity value for the full length, and the page now says so instead of showing a figure measured elsewhere.
ChangedMethod
The fidelity check now hangs on the amount of text that actually occurs, not on the selected step. A contract of 120,000 tokens needs the evidence for 128,000, not the one for a million; the same contract in German can cross the line, because the tokeniser makes more tokens of it. And when no model in the field carries an evidenced value for the full length, the question stops being a reason for exclusion and becomes a note for all of them.
AddedPrices
Fast mode now also applies to gpt-6-astra, though not together with European data residency. The other models in the tier stay as they were. In the process it emerged that the surcharge is not double everywhere: gpt-5.5 costs 2.5 times, older OpenAI models between 1.67 and 1.82 times. The calculator now uses the rate published per model instead of a flat factor.
Sources: OpenAI, Preisseite
FixedPrices
Peak hours at DeepSeek apply Monday to Friday only. That now stands at the model: 130 of the 168 hours in a week are off-peak at half the rate, the entire weekend included. The list price remains the peak rate.
Sources: DeepSeek, Preisseite
FixedMethod
Processing in the EU: on its own API Anthropic offers only the US or the global network, and an EU value does not exist there. The ten percent surcharge comes from Amazon Bedrock and Google Cloud and applies there from the 4.5 models on, including Claude Haiku 4.5, Opus 4.5 and Sonnet 4.5, which do not know the parameter themselves. Those three did not carry the surcharge before. Mistral has run its own EU and US endpoints since 11 August 2026, also at ten percent.
Sources: Anthropic, Datenresidenz · Mistral, API-Preise
ChangedAnchors
The writing anchor sits on the board as of 2 September 2026: top value 1504, threshold therefore 1484 instead of 1486. Gemini 3.8 Flash and Claude Fable 5.1 join, Claude Opus 5 and Gemini 3.1 Pro fall below the threshold. The research anchor sits on 8 September 2026 and newly lists gpt-6-astra at 91.5 and Kimi K3 at 91.2 percent. The Epoch capability index has been refitted and now covers 50 models instead of 40.
Sources: arena.ai, Creative Writing · BenchLM, BrowseComp · Epoch AI, Capabilities Index
FixedLiterature
Three sources stated more precisely: the price reversal paper is now cited in its second version of 28 May 2026, cost-of-pass in its second of 26 February 2026. On the 1.2 uplift to the attempt count, the page now separates evidence from assumption: Yang measures that assuming independent tries overstates the hit rate after three attempts by 17.4 points; the size of the uplift is not evidenced by that.
Sources: Chen et al., arXiv 2603.23971 · Erol et al., arXiv 2504.13359 · Yang, arXiv 2605.08563
DeprecatedModels
Claude Fable 5 is marked superseded, since Anthropic lists it under "Legacy models"; Mythos 5 follows it because 5.1 takes over the same role (the legacy list does not name Mythos, which runs through a separate programme). At the hosters the same applies to the lines whose successor is in the set: Qwen3.7-Max, Qwen3.6-Plus, Qwen3.5 and GLM-5.2.
Sources: Anthropic, Modellübersicht
ChangedRate
Rate moved to the ECB reference rate of 9 September 2026 (1.1652 USD per euro), previously 1.1576 of 18 August 2026.
Sources: EZB, Euro-Referenzkurs
AddedInterface
This change log. It shows what moved between two dates and where the figure comes from.
ChangedInterface
Title and description now name the counted number of models and the cost per solved task. They previously read "calculate cost per task", which named neither the scope nor the difference to other calculators.
ChangedInterface
The business case tiles lead with the problem instead of the reference: what is being calculated, what matters, and the evidence in the disclosure below. The markdown versions of both language editions for machines were added.
AddedMethod
The requirement profile: five controls for error cost, review, thinking effort, text volume and language. Since then the page calculates expected cost per solved task instead of cost per attempt, with a capability anchor per business case, a blocked zone for cases that must not go wrong, and a second opinion from the literature beside every result.
ChangedPrices
Since moving to the current versions, DeepSeek bills at peak and off-peak rates; the list price is the peak rate and sits three times above the old figure. Google sells Gemini 3.6 and 3.7 Flash at half price until the end of the year. Anthropic makes the introductory price of Sonnet 5 the standard, and the announced increase is dropped. OpenAI charges a surcharge for writing to the cache on the gpt-5.6 series.
FixedPrices
The long-context tiers of the entire gpt-5.6 series were missing; the long-document preset underestimated these models by up to half. Mistral has published a cache read price since August, so the line "Mistral lists no cache price" was out of date. The capability column from the Epoch index, attempts per model from the solve rate and five model classes were added.
ChangedModels
The set grew from 27 to 75 models. New: the second price tier for large inputs, the status per model and the context window. Since then the cache has been calculated with writes and reads at the rates published per model instead of a flat factor, and four pricing errors were fixed, among them a model that stood five times too expensive.
AddedInterface
First version of the calculator: 22 models, cost per task from input, output, cache and attempts, in euros.