Skip to main content
Comparison guide

Is the Cheapest AI Model Good Enough? What the Benchmarks Show

Winner

The cheap tier on recall and application; the frontier tier on novel reasoning

GPT-5.6 Sol, Terra and Luna; DeepSeek V4; Gemini; Grok

Best for

Deciding whether to move production traffic down a price tier

Research-based — not hands-on tested. Our scores are editorial judgements compiled from vendor documentation, published pricing and independent user reports. How we review.

Twenty-five times the price buys about two points. That is the finding, and it holds across most of what people actually use these models for — until you hit the one benchmark where the expensive model is forty-three times better and the cheap one scores essentially zero.

The comparison that is actually fair

Most cheap-versus-expensive model comparisons are built on sand, because they line up scores from different labs. Every vendor runs benchmarks on its own harness, with its own scaffolding, its own number of allowed attempts, and its own choice of which benchmarks to publish at all. Two models quoting 62 per cent on a benchmark with the same name may have been measured in meaningfully different ways.

There is one comparison that escapes this problem: comparing tiers inside a single vendor lineup. When OpenAI publishes GPT-5.6 Sol, Terra and Luna in one table, the same lab measured all three the same way on the same day. Whatever the harness does, it does it equally to each. The differences are real.

So that is where to start. Sol costs 5.00 dollars per million input tokens and 30.00 per million output. Luna costs 0.20 and 1.20. Twenty-five times the price, same lab, same table.

Benchmark

Published score, with the list price beside it

Read these across vendors with care. Every score here is published by the lab that made the model, on its own harness, with its own scaffolding — and each lab chooses which benchmarks to show at all. Comparing tiers inside one lineup is sound, because the same lab ran them the same way. Comparing one vendor against another is indicative, not decisive. There is deliberately no combined “intelligence score” on this page, because no such number is published by anyone.

On most work, the gap is small

On graduate-level science questions, Sol scores 94.6 and Luna scores 92.3. That is a gap of 2.3 points for twenty-five times the money.

On resolving real GitHub issues, Sol scores 64.6 and Luna 62.7 — a gap of 1.9 points. On multi-step command-line work, 88.8 against 84.7. On the coding-agent index, 80 against 74.6.

Four benchmarks, four gaps between two and six points. If your work looks like any of those — answering questions from a knowledge base, fixing well-specified bugs, running scripted multi-step tasks — the honest reading is that the cheap tier is doing most of the job, and the frontier tier is charging a very large premium for the last few points.

Whether those few points matter is a real question, not a rhetorical one. On a task where a 2 per cent error rate is annoying and recoverable, they probably do not. On a task where each failure costs a human twenty minutes of cleanup, two points across a hundred thousand requests is a lot of twenty minutes.

Then there is ARC-AGI-3

Switch the chart to novel reasoning and the picture inverts completely. Sol scores 7.78. Terra scores 0.8. Luna scores 0.18.

That is not a few points. Sol is roughly forty-three times better than Luna, and Terra — the middle tier, at ten times Luna price — is already down at 0.8.

ARC-AGI-3 is built specifically to resist memorisation. The puzzles are designed so that having seen the internet does not help you; you have to work out a rule you have never encountered from a handful of examples. Every model on the market scores badly on it in absolute terms. The interesting thing is not the absolute number, it is that the spread between tiers explodes precisely here, on the one benchmark that cannot be answered from pattern recall.

That is the most useful signal on this page, and it is worth stating plainly:

The cheap tier has very nearly caught the frontier at retrieving and applying things that are already known. It has not come close on working out something genuinely new.

The same models look fine on the older version of that test

Here is the complication, and it is the most interesting number on this page.

Run the same cheap models against ARC-AGI-2, the previous revision of that benchmark, and they do well. On the ARC Prize Foundation leaderboard — which is run independently of the labs, on one harness for everyone — GPT-5.6 Luna scores 59.6 per cent and DeepSeek V4 Flash scores 61.4 per cent. Claude Opus 4.5, a frontier model from the previous generation, scores 37.6 per cent. Gemini 3 Pro scores 31.1 per cent.

So on ARC-AGI-2 the cheap 2026 models comfortably beat the expensive 2025 models. On ARC-AGI-3 the same cheap models score under one per cent.

Two readings are available and they have very different consequences. Either ARC-AGI-3 is simply much harder, or ARC-AGI-2 has stopped measuring what it was built to measure — because once a benchmark has been public long enough, its puzzle types leak into training data and solving them stops requiring novel reasoning at all. Both are probably true to some degree, and neither can be settled from published scores.

What it does settle is a practical point. A benchmark result gets less informative the longer the benchmark has been public, and the freshest test is usually the one where the price ladder still shows a real gap. When you see a cheap model matching a frontier model on a well-known benchmark, the first question is not “has the cheap model caught up” — it is “how long has that test been on the internet”.

What that means for choosing a model

It suggests a rule that is more useful than a leaderboard position, because it is about the shape of your task rather than a single ranking.

The cheap tier is likely enough when the answer exists somewhere in the model training data or in the context you give it, and the job is to find it, apply it, and format it. Classification, extraction, summarising, answering from documents you supply, fixing bugs that resemble bugs the world has already fixed, following a runbook. This is the large majority of production traffic.

The frontier tier starts to earn its price when the task requires holding several unfamiliar constraints together and deriving something that is not in the training data. Genuinely novel debugging, architectural decisions with competing trade-offs, work where the first plausible answer is usually the wrong one. The tell is not that the task is hard for a person — it is that the task is unfamiliar.

Notice that this cuts against the usual advice. “Use the big model for important things” is not quite right. Importance is about consequences; capability is about novelty. A high-stakes task that is completely routine may be perfectly safe on the cheap tier, and a low-stakes task that is genuinely strange may embarrass it.

Is today’s cheap tier as good as last year’s flagship?

This is the question people actually mean when they ask whether a cheap model is good enough: not “is it good”, but “is it as good as the expensive thing I was told to buy a while ago”. For one specific pairing the answer is now clearly yes — and for a more recent pairing it is still clearly no. The line falls somewhere around February 2026.

All figures below were retrieved on 7 August 2026 from vendor model cards and named independent leaderboards. ARC-AGI-2 scores are quoted with the reasoning-effort tier they were measured at, because that setting moves the number more than the choice of model does.

GPT-5.6 Luna, max effort, vs… Launched ARC-AGI-2 Cost / task Result
Claude Opus 4.5 (Thinking, 64k) Nov 2025 37.6% vs Luna 59.6% $2.20 vs $0.177 Luna wins by 22.0pp, at about a twelfth of the cost
Gemini 3 Pro Nov 2025 31.11% vs Luna 59.6% $0.81 vs $0.177 Luna wins by 28.5pp
Claude Opus 4.6 Feb 2026 64.6–69.2% across effort tiers Luna loses
GPT-5.4 Mar 2026 67.5–83.3% at High / XHigh / Pro-XHigh Luna loses clearly
ARC-AGI-2 scores and per-task costs from the ARC Prize leaderboard (arcprize.org), retrieved 7 August 2026.

The nine-month line

Read down that table and the pattern is not “cheap models have caught up”. It is that cheap models have caught up with a specific vintage. GPT-5.6 Luna beats the flagships of November 2025 and loses to the flagships of February and March 2026. If you are asking whether today’s cheapest tier replaces the expensive model you were using, the honest answer depends entirely on when you bought it — roughly nine months of progress, not three.

Two other measures, both less flattering

ARC-AGI-2 is the benchmark where the cheap tier looks best, so it should not be read alone.

  • On the Artificial Analysis Intelligence Index v4.1.1, Luna scores 52 against Claude Opus 4.5’s 36, Claude Opus 4.6’s 39 and Gemini 3 Pro’s 41 — but GPT-5.4 scores 53. Every cheap-tier model in this set still trails GPT-5.4 on that index, Luna by a single point.
  • On LMArena / Arena.ai text Elo, Luna sits at 1450 against GPT-5.4’s 1465 and Claude Opus 4.5’s 1469. It trails both, though 15–19 Elo is a modest gap.
  • On SWE-bench Verified run under one standardised harness (swebench.com, mini-SWE-agent v2), the only cheap-tier model on the board is Claude Haiku 4.5 at 66.60%, trailing Opus 4.5’s 76.80% by 10.2 points. Luna has no entry there, so no same-harness coding comparison is possible for it.

What actually changed is the price

The capability story is mixed. The pricing story is not. Luna at its 30 July 2026 price of $0.20 in / $1.20 out per million tokens is roughly 25× cheaper on input and 21× cheaper on output than Claude Opus 4.5’s launch price of $5 / $25 — while scoring higher on ARC-AGI-2 and close behind on Elo. DeepSeek V4 Flash at $0.14 / $0.28 is cheaper still.

So the useful framing is not “cheap models are as good now”. It is that the price of roughly-November-2025 frontier capability has collapsed, while the actual frontier moved on. If your workload was well served by a late-2025 flagship, you can probably run it for a fraction of the cost today. If you were relying on a 2026 flagship, nothing here says you can drop down.

Reading across vendors, carefully

Once you leave a single vendor lineup, the ground gets softer. DeepSeek V4 Flash publishes 88.1 on the same named science benchmark, at 0.14 dollars per million input tokens. That is a striking number next to Sol 94.6 at 5.00. But DeepSeek ran that evaluation, not OpenAI, and the two labs did not agree a common method.

The gap between vendors is also partly a gap in what they choose to publish. DeepSeek publishes SWE-bench Verified; OpenAI publishes SWE-Bench Pro. Those are different benchmarks with different score scales, and putting them in the same column would be a mistake even though the names look similar. Where a model is missing from a chart on this page, it is because that vendor did not publish that benchmark — not because we estimated a zero.

Treat cross-vendor numbers as a reason to run your own test, not as a result.

How to check it on your own workload

A benchmark is a proxy. Your evaluation set is not, and building one is a smaller job than it sounds.

  1. Collect twenty to fifty real inputs from your actual traffic. Not invented examples — real ones, including the messy inputs you would rather forget about.
  2. Write down what a correct output looks like in checkable terms. The JSON parses and has the required keys. The total matches the invoice. The summary names every party. If you cannot state the check, you cannot run the test.
  3. Run the cheapest candidate first. Starting at the bottom tells you what you actually need; starting at the top only tells you that the expensive model works, which you already assumed.
  4. Look at the failures, not the score. A model that fails 5 per cent of cases randomly is a different proposition from one that fails 5 per cent of cases all in the same category. The second is fixable with routing; the first is not.
  5. Re-run it when prices move. They move often. OpenAI cut two of its three tiers on 30 July 2026, one of them by 80 per cent.

The part nobody can benchmark for you

Everything above is about published numbers, and published numbers have a ceiling as evidence. They are self-reported, they are chosen by the vendor, the harnesses differ, and benchmark contamination is a real and largely unmeasured problem — a model that has seen a benchmark test set during training will score well on it without being better at anything.

What the numbers can do is tell you where to look. They say: the cheap tier is close on recall and application, far behind on novelty, and the price difference is large enough that the question deserves a real answer rather than a default.

The rest is a twenty-input evaluation set and an afternoon.