AI Model Comparison 2026: Cost vs What Is Good Enough
Depends on your token mix - the dashboard ranks for your own volumes
Claude, GPT-5.6, Gemini, DeepSeek and Grok API models
Teams costing an API workload before they commit to a vendor
DeepSeek V4 Flash is cheapest on both the input and the output axis at…
Research-based — not hands-on tested. Our scores are editorial judgements compiled from vendor documentation, published pricing and independent user reports. How we review.
Twelve frontier models from five vendors, priced from pages the vendors publish themselves, with the arithmetic done for your actual workload rather than for a tidy round million tokens. Move the sliders and watch the bill move with them.
Set the workload and the whole table re-prices. Output tokens cost two to five times what input tokens cost, so the size of your bill depends as much on your input-to-output ratio as on which model you pick. At the prices published today the gap between the cheapest and dearest option here runs from about seventy times to about a hundred and seventy times, depending on where you put the dials.
Every published figure, with its source
| Model | Input$/1M | Output$/1M | Cached in$/1M | Context | Cost / month |
|---|
Prices read from each vendor's own published pricing page on 7 August 2026. A dash means the vendor does not publish that figure — we leave it empty rather than estimating it. Model prices move fast: OpenAI repriced the GPT‑5.6 tiers on 30 July 2026, eight days before this table was compiled. Check the vendor page before committing to a budget.
The spread is the story
At a modest support-chat volume, the gap between the cheapest and the most expensive model in this table is roughly a hundredfold. That is not a rounding difference you can absorb. It is the difference between a line item nobody notices and a line item that gets a meeting.
Two things drive it, and only one of them is obvious.
The obvious one is tier. Every vendor now ships a cheap model, a workhorse, and a frontier model, and the price ladder inside a single vendor lineup can be steeper than the gap between vendors. OpenAI repriced GPT-5.6 on 30 July 2026 and cut Luna by 80 per cent while leaving Sol untouched, which stretched the distance between the top and bottom of one product family from roughly five times to roughly twenty-five.
The less obvious one is your input-to-output ratio. Output tokens cost between two and five times what input tokens cost on every model in this table, so a job that reads a lot and writes almost nothing behaves nothing like one that reads a paragraph and writes a page.
Worth being precise about what that does and does not change, because it is easy to overclaim. At the prices published today it does not change who wins: DeepSeek V4 Flash is the cheapest model on both the input and the output axis, so it leads at every mix on the dial. What the mix does change is the order in the middle of the table, where GPT-5.6 Luna and DeepSeek V4 Pro trade places once output starts to dominate, and it changes the size of the gap enormously, from roughly seventy times on an input-heavy extraction job to roughly a hundred and seventy times on an output-heavy drafting one.
That second number is the one worth acting on. The question is rarely which model is cheapest in the abstract; it is how much you would actually save by moving, and whether that saving is worth the work of re-testing.
What good enough actually means
The honest answer to which model is good enough is that nobody can tell you from a price table, and you should be suspicious of any page that pretends otherwise.
Price is a published, directly comparable number. Every vendor states it in the same units, and a dollar is a dollar. Capability is not like that. Vendors publish their own benchmark scores, run on their own harnesses, chosen by their own marketing teams, and those numbers are not comparable across vendors in the way prices are. Two models quoting the same score on the same named benchmark may have been evaluated under different prompting, different attempt budgets, and different scaffolding.
So the useful question is not which model is best. It is: what is the cheapest model that clears the bar for my specific task, and how would I know if it stopped clearing it?
That question has a method, and the method is the same regardless of which models are in fashion:
- Write down what failure looks like before you test anything. Not a good answer, but a checkable one: the JSON parses, the number matches the invoice, the summary names all three parties.
- Start at the bottom of the price ladder, not the top. The cheap tier is where most production traffic belongs, and starting cheap tells you what you actually need rather than what you assumed.
- Move up one tier only when a specific, repeatable failure forces you to. One bad output is noise. A failure mode you can reproduce is a reason.
- Re-test when prices move. They move often, and an 80 per cent cut can make a model you rejected on cost the obvious default overnight.
The cached-input column is the one most people miss
Look at the cached-input column. Every vendor in this table except Anthropic publishes one, and the discounts are not small: DeepSeek V4 Flash charges 0.0028 dollars per million cached tokens against 0.14 for a cache miss, a fiftyfold difference. OpenAI, Google and xAI all price cached input at roughly one tenth of their standard input rate.
If your application resends a large fixed prefix on every request, which describes most agents, most retrieval pipelines and most classification jobs, then the cached rate is closer to your real input cost than the headline number is. A model that looks mid-priced on the sticker can be the cheapest thing on the page once the prefix is cached, and the dashboard above deliberately shows both so the comparison is not flattering to the wrong model.
The catch is that caching is not free or automatic. Each vendor sets a minimum cacheable prefix length and a time-to-live, and a prompt that changes even slightly at the front invalidates everything after it. A timestamp at the top of a system prompt is enough to reduce the cache hit rate to zero.
Two models we checked and left out
A comparison table is only as trustworthy as the things it refuses to include, so here is what did not make it and why.
Meta Llama. There is no live first-party API pricing page. The Llama API was wound down in July 2026, and the only prices available now belong to third-party hosts running the weights on their own hardware. Those are real prices, but they are the price of a host, not the price of a model, and putting them in the same column as a vendor rate would compare two different things. Llama is omitted rather than approximated.
Mistral Large 3. Two Mistral pages give two different prices. The general pricing page states 2.00 dollars in and 6.00 out per million tokens; the dedicated API pricing table, the one tied to the exact model identifier, states 0.50 and 1.50. That is a fourfold discrepancy inside one vendor site, and we could not establish which page is stale. Rather than pick one and present it as settled, the model is left out and the contradiction is reported here. If you are costing Mistral, read both pages and get the answer in writing.
What this page deliberately does not do
It does not give you a single intelligence score. No such number exists. Every vendor publishes its own benchmark results, on its own harness, with its own scaffolding and its own choice of which benchmarks to show. Averaging those into one figure would produce something that looks authoritative and means nothing, and reasonable-looking composite scores are exactly how comparison pages mislead people.
It also does not rank models by quality. We have not run a controlled evaluation of these models against each other, so we do not claim one. Where a vendor publishes a benchmark result we can cite it as what the vendor says; that is a different claim from what we found.
What the page does give you is the part that is genuinely comparable: published prices, published context windows, and the arithmetic for your workload, with a source and a retrieval date on every number.
Sources and method
Every figure comes from a page published by the vendor itself. Aggregator and comparison sites were used only to locate vendor pages, never as the source of a number. Where a vendor does not publish something, the cell is empty and stays empty.
Prices were read on 7 August 2026. They will go out of date, and faster than you expect: OpenAI cut two of its three tiers on 30 July, Google replaced Gemini 3.5 Flash on 21 July, and the DeepSeek pricing page currently carries a standing warning that a significant increase is planned. Treat this table as a starting point and the vendor page as the authority.