Skip to main content
Comparison guide

AI Model Comparison 2026: Cost vs What Is Good Enough

Winner

Depends on your token mix - the dashboard ranks for your own volumes

Claude, GPT-5.6, Gemini, DeepSeek and Grok API models

Best for

Teams costing an API workload before they commit to a vendor

Pricing

The callout above names today's cheapest model for your chosen workloa…

Research-based — not hands-on tested. Our scores are editorial judgements compiled from vendor documentation, published pricing and independent user reports. How we review.

Comparing against a model you used before

If you were using an earlier flagship a few months ago, here is how it stacks up against what is current today.

How this compares with an earlier flagship

Claude Opus 4.6 — then (5 Feb 2026) $1,175
Claude Opus 4.6 — at today's price $1,175
Best alternative today No alternative has enough comparable evidence yet.

Cheapest listed price: $57 for DeepSeek-V4.1-Flash (comparable evidence not yet confirmed — not a recommendation).

Not enough comparable evidence to summarize this workload yet.

These are API model comparisons. An agent product (Claude Code, Codex, Gemini CLI, Cursor…) adds tools, memory, orchestration and an execution environment; a model benchmark does not establish what an agent application can do.

Current API models from the major vendors, priced from pages the vendors publish themselves, with the arithmetic done for your actual workload rather than for a tidy round million tokens. Move the sliders and watch the bill move with them.

Set the workload and the whole table re-prices. Output tokens typically cost several times what input tokens cost, so the size of your bill depends as much on your input-to-output ratio as on which model you pick. The callout below states today's actual spread between the cheapest and dearest option for the workload you choose.

Workload
Vendors
Estimated monthly API cost at the selected volumes

Every published figure, with its source

Model Input$/1M Output$/1M Cached in$/1M Context Cost / month

Prices read from each vendor's own published page between 11 September 2026 and 30 September 2026. A dash means the vendor does not publish that figure — we leave it empty rather than estimating it. Model prices move fast and change without notice; check each vendor's own pricing page before committing to a budget.

The spread is the story

At a modest support-chat volume, the gap between the cheapest and the most expensive model in this table is large enough to decide a budget on its own. The callout under the chart gives the exact multiple for the workload you set.

Two things drive it, and only one of them is obvious.

The obvious one is tier. Every vendor now ships a cheap model, a workhorse, and a frontier model, and the price ladder inside a single vendor lineup can be steeper than the gap between vendors. OpenAI repriced GPT-5.6 on 30 July 2026 and cut Luna by 80 per cent while leaving Sol untouched, which stretched the distance between the top and bottom of one product family from roughly five times to roughly twenty-five.

The less obvious one is your input-to-output ratio. Output tokens cost between about three and six times what input tokens cost on the models in this table, so a job that reads a lot and writes almost nothing behaves nothing like one that reads a paragraph and writes a page.

Worth being precise about what the mix does and does not change. When one model is cheapest on both the input and the output price, it leads at every setting of the dial, and the mix only reorders the middle of the table and changes the size of the gap. The callout names today’s leader and that gap for the mix you choose.

That second number is the one worth acting on. The question is rarely which model is cheapest in the abstract; it is how much you would actually save by moving, and whether that saving is worth the work of re-testing.

What good enough actually means

The honest answer to which model is good enough is that nobody can tell you from a price table, and you should be suspicious of any page that pretends otherwise.

Price is a published, directly comparable number. Every vendor states it in the same units, and a dollar is a dollar. Capability is not like that. Vendors publish their own benchmark scores, run on their own harnesses, chosen by their own marketing teams, and those numbers are not comparable across vendors in the way prices are. Two models quoting the same score on the same named benchmark may have been evaluated under different prompting, different attempt budgets, and different scaffolding.

So the useful question is not which model is best. It is: what is the cheapest model that clears the bar for my specific task, and how would I know if it stopped clearing it?

That question has a method, and the method is the same regardless of which models are in fashion:

  1. Write down what failure looks like before you test anything. Not a good answer, but a checkable one: the JSON parses, the number matches the invoice, the summary names all three parties.
  2. Start at the bottom of the price ladder, not the top. The cheap tier is where most production traffic belongs, and starting cheap tells you what you actually need rather than what you assumed.
  3. Move up one tier only when a specific, repeatable failure forces you to. One bad output is noise. A failure mode you can reproduce is a reason.
  4. Re-test when prices move. They move often, and an 80 per cent cut can make a model you rejected on cost the obvious default overnight.

The cached-input column is the one most people miss

Look at the cached-input column. Every vendor in this table publishes one, and the discounts are often large, from about a quarter of the standard input rate down to a fiftieth. The exact figure for each model is in the table.

If your application resends a large fixed prefix on every request, which describes most agents, most retrieval pipelines and most classification jobs, then the cached rate is closer to your real input cost than the headline number is. A model that looks mid-priced on the sticker can be the cheapest thing on the page once the prefix is cached, and the dashboard above deliberately shows both so the comparison is not flattering to the wrong model.

The catch is that caching is not free or automatic. Each vendor sets a minimum cacheable prefix length and a time-to-live, and a prompt that changes even slightly at the front invalidates everything after it. A timestamp at the top of a system prompt is enough to reduce the cache hit rate to zero.

Two models we checked and left out

A comparison table is only as trustworthy as the things it refuses to include, so here is what did not make it and why.

Meta Llama. There is no live first-party API pricing page. The Llama API was wound down in July 2026, and the only prices available now belong to third-party hosts running the weights on their own hardware. Those are real prices, but they are the price of a host, not the price of a model, and putting them in the same column as a vendor rate would compare two different things. Llama is omitted rather than approximated.

Mistral Large 3. On 7 August 2026 Mistral’s two pricing pages disagreed fourfold ($2.00 and $6.00 against $0.50 and $1.50 per million tokens). On 11 September 2026 both pages showed $0.50 and $1.50. We are keeping it out of this table until that figure has held through a full refresh cycle.

What this page deliberately does not do

It does not give you a single intelligence score. No such number exists. Every vendor publishes its own benchmark results, on its own harness, with its own scaffolding and its own choice of which benchmarks to show. Averaging those into one figure would produce something that looks authoritative and means nothing, and reasonable-looking composite scores are exactly how comparison pages mislead people.

It also does not rank models by quality. We have not run a controlled evaluation of these models against each other, so we do not claim one. Where a vendor publishes a benchmark result we can cite it as what the vendor says; that is a different claim from what we found.

What the page does give you is the part that is genuinely comparable: published prices, published context windows, and the arithmetic for your workload, with a source and a retrieval date on every number.

Sources and method

Every figure comes from a page published by the vendor itself. Aggregator and comparison sites were used only to locate vendor pages, never as the source of a number. Where a vendor does not publish something, the cell is empty and stays empty.

Each price carries the date it was read from the vendor page. Prices go out of date faster than you expect: OpenAI cut two of its three GPT-5.6 tiers on 30 July 2026 and temporarily cut Sol on 21 August, Google moved its Flash models to introductory prices that end on 31 December 2026, and DeepSeek now charges different rates at peak and off-peak hours. Treat this table as a starting point and the vendor page as the authority.