Skip to main content
Interactive comparison

AI Model Benchmarks & Real Cost per Solved Task

What AI models score, and what those runs actually cost — not just today's API list price.

Compare benchmark performance and cost

Human-filtered subset of SWE-bench; resolves real GitHub issues end-to-end. About this benchmark.

Showing 7 results evaluated under the SAME settings (mini-SWE-agent 2.0.0, reasoning effort: high) — the only results on this page that are directly comparable to each other.

Benchmark score and cost for SWE-bench Verified
Current API price (in / out per 1M) Evaluated Source
Claude Opus 4.5, mini-SWE-agent 2.0.0, reasoning effort: high 76.8% $0.75 $0.98 $0.01 $5.00 / $25 17 Feb 2026 SWE-bench (Princeton/Stanford team)
Gemini 3 Flash (Preview), mini-SWE-agent 2.0.0, reasoning effort: high 75.8% $0.36 $0.47 $0.00 not published 17 Feb 2026 SWE-bench (Princeton/Stanford team)
MiniMax M2.5, mini-SWE-agent 2.0.0, reasoning effort: high 75.8% $0.07 $0.10 $0.00 $0.300 / $1.20 17 Feb 2026 SWE-bench (Princeton/Stanford team)
GPT-5.2, mini-SWE-agent 2.0.0, reasoning effort: high 72.8% $0.47 $0.65 $0.01 $1.75 / $14 17 Feb 2026 SWE-bench (Princeton/Stanford team)
Claude Sonnet 4.5, mini-SWE-agent 2.0.0, reasoning effort: high 71.4% $0.66 $0.92 $0.01 $3.00 / $15 17 Feb 2026 SWE-bench (Princeton/Stanford team)
Gemini 3 Pro (Preview), mini-SWE-agent 2.0.0, reasoning effort: high 69.6% $0.96 $1.38 $0.01 $2.00 / $12 26 Feb 2026 SWE-bench (Princeton/Stanford team)
Claude Haiku 4.5, mini-SWE-agent 2.0.0, reasoning effort: high 66.6% $0.33 $0.50 $0.00 $1.00 / $5.00 17 Feb 2026 SWE-bench (Princeton/Stanford team)

Cost / point is a descriptive ratio within this benchmark and cost basis only — not the price of gaining another point, and never comparable across benchmarks or against a non-percent score.

25 catalogued models are not shown for this benchmark: 22 model not mapped to a tracked model id in our catalogue; 3 the source records no cost for this run. See the full catalogue.

Other reported results (different evaluation settings) — 15 results across 10 other configurations

Not proven equivalent to the group above — different agent, agent version, or reasoning effort. Each row names its own configuration.

Benchmark score and cost for SWE-bench Verified
Current API price (in / out per 1M) Evaluated Source
Claude Opus 4.6, mini-SWE-agent 2.0.0, reasoning effort not stated by the source 75.6% $0.55 $0.73 $0.01 $5.00 / $25 17 Feb 2026 SWE-bench (Princeton/Stanford team)
Claude Opus 4.5, mini-SWE-agent 1.16.0, reasoning effort: medium 74.4% $0.72 $0.97 $0.01 $5.00 / $25 24 Nov 2025 SWE-bench (Princeton/Stanford team)
Gemini 3 Pro (Preview), mini-SWE-agent 1.15.0, reasoning effort not stated by the source 74.2% $0.46 $0.62 $0.01 $2.00 / $12 18 Nov 2025 SWE-bench (Princeton/Stanford team)
GPT-5.2, mini-SWE-agent 1.17.2, reasoning effort: high 71.8% $0.52 $0.72 $0.01 $1.75 / $14 11 Dec 2025 SWE-bench (Princeton/Stanford team)
Claude Sonnet 4.5, mini-SWE-agent 1.13.3, reasoning effort not stated by the source 70.6% $0.56 $0.79 $0.01 $3.00 / $15 29 Sep 2025 SWE-bench (Princeton/Stanford team)
GPT-5.2, mini-SWE-agent 1.17.2, reasoning effort not stated by the source 69.0% $0.27 $0.39 $0.00 $1.75 / $14 11 Dec 2025 SWE-bench (Princeton/Stanford team)
GPT-5.1, mini-SWE-agent 1.15.0, reasoning effort: medium 66.0% $0.31 $0.46 $0.00 $1.25 / $10 20 Nov 2025 SWE-bench (Princeton/Stanford team)
GPT-5, mini-SWE-agent 1.7.0, reasoning effort: medium 65.0% $0.28 $0.43 $0.00 $1.25 / $10 7 Aug 2025 SWE-bench (Princeton/Stanford team)
GPT-5 mini, mini-SWE-agent 1.7.0, reasoning effort: medium 59.8% $0.04 $0.06 $0.00 $0.250 / $2.00 7 Aug 2025 SWE-bench (Princeton/Stanford team)
o3, mini-SWE-agent 1.0.0, reasoning effort not stated by the source 58.4% $0.33 $0.57 $0.01 $2.00 / $8.00 26 Jul 2025 SWE-bench (Princeton/Stanford team)
GPT-5 mini, mini-SWE-agent 2.0.0, reasoning effort not stated by the source 56.2% $0.05 $0.08 $0.00 $0.250 / $2.00 17 Feb 2026 SWE-bench (Princeton/Stanford team)
GLM-4.6, mini-SWE-agent 1.17.1, reasoning effort not stated by the source 55.4% $0.10 $0.17 $0.00 $0.600 / $2.20 1 Dec 2025 SWE-bench (Princeton/Stanford team)
Gemini 2.5 Pro, mini-SWE-agent 1.0.0, reasoning effort not stated by the source 53.6% $0.29 $0.54 $0.01 $1.25 / $10 26 Jul 2025 SWE-bench (Princeton/Stanford team)
GPT-5 nano, mini-SWE-agent 1.7.0, reasoning effort: medium 34.8% $0.04 $0.11 $0.00 $0.050 / $0.400 7 Aug 2025 SWE-bench (Princeton/Stanford team)
Gemini 2.5 Flash, mini-SWE-agent 1.0.0, reasoning effort not stated by the source 28.7% $0.13 $0.47 $0.00 $0.300 / $2.50 26 Jul 2025 SWE-bench (Princeton/Stanford team)

Cost / point is a descriptive ratio within this benchmark and cost basis only — not the price of gaining another point, and never comparable across benchmarks or against a non-percent score.

Whole-workload equivalence analysis (optional) Pick a model you used before to see whether anything newer is confirmed to match it on every material capability for your own workload — a stricter, narrower question than the benchmark comparison above.
Claude Opus 4.6 was released 5 Feb 2026. Newer Anthropic flagships by then: Claude Opus 4.7, Claude Opus 4.8, Claude Fable 5.
No confirmed match today

$1,233 Claude Opus 4.6 — current cost for this workload

GPT-5.6 Sol has the most comparable evidence of any alternative, but is below baseline on coding, which is 55% of your workload.

Full baseline match

No model has enough comparable evidence to be called a baseline match for this workload. Closest candidates: Gemini 3 Flash (Preview) (55% covered), Kimi K3 (30% covered), Claude Fable 5.1 (30% covered).

Workload & assumptions Coding · 5K tasks/month · Standard
No confirmed match today$1,233 baseline/mo
Start with a preset

Each slider is a relative weight across your workload; the percentage shown is that weight's share of the total once every slider is combined. This is your workload mix — a separate number from a capability explanation like "55% of your workload" below, which names that capability's own internal weight WITHIN one category (e.g. how much of "coding" work is judged on the coding capability itself vs. tool use), not a second mix you set.

100% of your work
0% of your work
0% of your work
0% of your work
0% of your work
Task complexity
5K tasks / month
Priorities (0 = ignore, 3 = most important)

These reorder the full "All models" list only — they never change which model is picked as best value, strongest, or cheapest meeting the baseline, and never change a capability verdict.

ignore
ignore
ignore
ignore
Advanced assumptions

These per-task token counts, cache and retry assumptions are illustrative defaults, not measurements — editable here, applied only to your estimate.

See why, for each role
Strongest supported option
GPT-5.6 Sol has the most comparable evidence of any alternative, but is below baseline on coding, which is 55% of your workload.
Cheapest meeting the baseline
No model has enough comparable evidence to be called a baseline match for this workload. Closest candidates: Gemini 3 Flash (Preview) (55% covered), Kimi K3 (30% covered), Claude Fable 5.1 (30% covered).
Best value
No model qualifies: Gemini 3 Flash (Preview) has comparable evidence for only 55% of your workload: there is no comparable result for tool use and autonomous execution and reasoning.

No model available today is confirmed to match Claude Opus 4.6 on your selected workload. GPT-5.6 Sol has the most comparable evidence of any alternative, but is below baseline on coding, which is 55% of your workload.

These are API model comparisons. An agent product (Claude Code, Codex, Gemini CLI, Cursor…) adds tools, memory, orchestration and an execution environment; a model benchmark does not establish what an agent application can do.

Claude Opus 4.6 — then (13 Jun 2026) $1,233
Claude Opus 4.6 — at today's price $1,233
Best value alternative today No alternative has enough comparable evidence yet.

Prices checked between 11 Sep 2026 and 12 Sep 2026 · dataset 2026-09-12.14

Research-based: we compare published prices and published benchmark results. We have not run these models ourselves, and there is no overall score.

Capability matrix

Model Coding Tool use Autonomous execution Reasoning
Claude Fable 5.1
Kimi K3
Gemini 3.5 Flash
GPT-6 Astra
Gemini 3.8 Flash
Claude Sonnet 4.6
Claude Opus 4.7
Claude Opus 4.8
Claude Fable 5
Grok 4
Grok 4 Fast
Grok 4.1 Fast
Grok 4.3
Grok 4.5
Grok 4.6
GPT-5.4
GPT-5.4 mini
GPT-5.4 nano
GPT-5.5
GPT-5.5 Pro
GPT-5.6 Luna
o3-pro
o3-mini
o3 deep research
DeepSeek-V4-Pro
Gemini 3.6 Flash
Gemini 3 Flash (Preview)
Gemini 2.5 Flash-Lite
Qwen3.8-Max
Kimi K2.6
Mistral Large 3
MiniMax M2.5
MiniMax M2.7
Claude Mythos 5
Claude Mythos 5.1
DeepSeek-V4.1-Flash
Claude Opus 5
GPT-5.6 Sol
GPT-5.6 Terra
Gemini 3.7 Flash
Gemini 3.1 Pro (Preview)
Claude Sonnet 4.5
Claude Sonnet 5
GPT-5
GPT-5 mini
GPT-5 nano
GPT-5.1
o3
Gemini 2.5 Pro
Gemini 2.5 Flash
GLM-4.6
GLM-5.3
Claude Haiku 4.5
GPT-5.2
Claude Opus 4.5

Timeline

Model Released Status Input $/1M Output $/1M
Claude Fable 5Anthropic 9 Jun 2026 Generally available $10 $50
Claude Fable 5.1Anthropic 1 Sep 2026 Generally available $10 $50
Claude Haiku 4.5Anthropic 15 Oct 2025 Generally available $1.00 $5.00
Claude Mythos 5Anthropic 9 Jun 2026 Preview $10 $50
Claude Mythos 5.1Anthropic 1 Sep 2026 Preview $10 $50
Claude Opus 4.5Anthropic 24 Nov 2025 Generally available $5.00 $25
Claude Opus 4.6Anthropic 5 Feb 2026 Generally available $5.00 $25
Claude Opus 4.7Anthropic 16 Apr 2026 Generally available $5.00 $25
Claude Opus 4.8Anthropic 28 May 2026 Generally available $5.00 $25
Claude Opus 5Anthropic 24 Jul 2026 Generally available $5.00 $25
Claude Sonnet 4.5Anthropic 29 Sep 2025 Generally available $3.00 $15
Claude Sonnet 4.6Anthropic 17 Feb 2026 Generally available $3.00 $15
Claude Sonnet 5Anthropic 30 Jun 2026 Generally available $2.00 $10
DeepSeek-V4-ProDeepSeek 13 Aug 2026 Generally available $1.32 $3.96
DeepSeek-V4.1-FlashDeepSeek 10 Sep 2026 Generally available $0.300 $1.20
GLM-4.6Z.ai (GLM) Release date not published; listed from 11 Sep 2026, when its price was first read. Generally available $0.600 $2.20
GLM-5.3Z.ai (GLM) 14 Aug 2026 Generally available $1.40 $4.40
GPT-5OpenAI 7 Aug 2025 Deprecated Deprecated $1.25 $10
GPT-5 miniOpenAI 7 Aug 2025 Deprecated Deprecated $0.250 $2.00
GPT-5 nanoOpenAI 7 Aug 2025 Deprecated Deprecated $0.050 $0.400
GPT-5.1OpenAI 13 Nov 2025 Deprecated Deprecated $1.25 $10
GPT-5.2OpenAI 11 Dec 2025 Deprecated Deprecated $1.75 $14
GPT-5.4OpenAI 5 Mar 2026 Generally available $2.50 $15
GPT-5.4 miniOpenAI 17 Mar 2026 Generally available $0.750 $4.50
GPT-5.4 nanoOpenAI 17 Mar 2026 Generally available $0.200 $1.25
GPT-5.5OpenAI 24 Apr 2026 Generally available $5.00 $30
GPT-5.5 ProOpenAI 24 Apr 2026 Generally available $30 $180
GPT-5.6 LunaOpenAI 9 Jul 2026 Generally available $0.200 $1.20
GPT-5.6 SolOpenAI 9 Jul 2026 Generally available $4.00 $20
GPT-5.6 TerraOpenAI 9 Jul 2026 Generally available $2.00 $12
GPT-6 AstraOpenAI 3 Sep 2026 Generally available $10 $50
Gemini 2.5 FlashGoogle 17 Jun 2025 Generally available $0.300 $2.50
Gemini 2.5 Flash-LiteGoogle 22 Jul 2025 Generally available $0.100 $0.400
Gemini 2.5 ProGoogle 17 Jun 2025 Generally available $1.25 $10
Gemini 3 Flash (Preview)Google 17 Dec 2025 Preview
Gemini 3.1 Pro (Preview)Google 19 Feb 2026 Preview $2.00 $12
Gemini 3.5 FlashGoogle 19 May 2026 Generally available $1.50 $9.00
Gemini 3.6 FlashGoogle 21 Jul 2026 Generally available $0.750 $3.75
Gemini 3.7 FlashGoogle 13 Aug 2026 Generally available $0.750 $3.75
Gemini 3.8 FlashGoogle 2 Sep 2026 Generally available $0.750 $3.75
Grok 4xAI 9 Jul 2025 Generally available
Grok 4 FastxAI 19 Sep 2025 Generally available $0.200 $0.500
Grok 4.1 FastxAI 19 Nov 2025 Generally available $0.200 $0.500
Grok 4.3xAI Release date not published; listed from 11 Sep 2026, when its price was first read. Generally available $1.25 $2.50
Grok 4.5xAI 16 Jul 2026 Generally available $2.00 $6.00
Grok 4.6xAI 12 Aug 2026 Generally available $2.00 $6.00
Kimi K2.6Moonshot AI (Kimi) Release date not published; listed from 11 Sep 2026, when its price was first read. Generally available $0.950 $4.00
Kimi K3Moonshot AI (Kimi) Release date not published; listed from 11 Sep 2026, when its price was first read. Generally available $3.00 $15
MiniMax M2.5MiniMax Release date not published; listed from 11 Sep 2026, when its price was first read. Deprecated Deprecated $0.300 $1.20
MiniMax M2.7MiniMax Release date not published; listed from 11 Sep 2026, when its price was first read. Generally available $0.300 $1.20
Mistral Large 3Mistral AI 1 Dec 2025 Generally available $0.500 $1.50
Qwen3.8-MaxAlibaba (Qwen) 3 Aug 2026 Generally available
o3OpenAI 16 Apr 2025 Deprecated Deprecated $2.00 $8.00
o3 deep researchOpenAI 26 Jun 2025 Deprecated Deprecated
o3-miniOpenAI 31 Jan 2025 Deprecated Deprecated $1.10 $4.40
o3-proOpenAI 10 Jun 2025 Deprecated Deprecated $20 $80

Catch-up by provider

  • Anthropic first met the baseline with Claude Opus 5 on 24 Jul 2026 — Above baseline.
  • OpenAI first met the baseline with GPT-5.5 on 24 Apr 2026 — Similar range.
  • Google first met the baseline with Gemini 3.1 Pro (Preview) on 19 Feb 2026 — Similar range. First cheaper than the baseline while meeting it: Gemini 3.1 Pro (Preview), 19 Feb 2026.
  • xAI not yet shown on comparable evidence.
  • DeepSeek not yet shown on comparable evidence.
  • Alibaba (Qwen) not yet shown on comparable evidence.
  • Moonshot AI (Kimi) not yet shown on comparable evidence.
  • Z.ai (GLM) not yet shown on comparable evidence.
  • Mistral AI not yet shown on comparable evidence.
  • MiniMax not yet shown on comparable evidence.
  • Meta not yet shown on comparable evidence.

No provider has yet matched the whole workload on comparable evidence.

Every option, ranked

Sorted by capability match against your baseline (the default ranking) — set a priority below to reorder it.

Meets baseline Released Status Savings Speed More details
Claude Fable 5.1Anthropic $2,395 Cannot confirm 30% 1 Sep 2026 Generally available -94% 67 tok/s
Release · status · savings · speed
Released
1 Sep 2026
Status
Generally available
Savings
-94%
Speed
67 tok/s
Kimi K3Moonshot AI (Kimi) $723 Cannot confirm 30% Release date not published; listed from 11 Sep 2026, when its price was first read. Generally available 41% 39 tok/s
Release · status · savings · speed
Released
— (Release date not published; listed from 11 Sep 2026, when its price was first read.)
Status
Generally available
Savings
41%
Speed
39 tok/s
Gemini 3.5 FlashGoogle $411 Cannot confirm 20% 19 May 2026 Generally available 67% no data
Release · status · savings · speed
Released
19 May 2026
Status
Generally available
Savings
67%
Speed
no data
GPT-6 AstraOpenAI $2,466 Cannot confirm 10% 3 Sep 2026 Generally available -100% 54 tok/s
Release · status · savings · speed
Released
3 Sep 2026
Status
Generally available
Savings
-100%
Speed
54 tok/s
Gemini 3.8 FlashGoogle $181 Cannot confirm 10% 2 Sep 2026 Generally available 85% 261 tok/s
Release · status · savings · speed
Released
2 Sep 2026
Status
Generally available
Savings
85%
Speed
261 tok/s
Claude Sonnet 4.6Anthropic $740 Cannot confirm 0% 17 Feb 2026 Generally available 40% no data
Release · status · savings · speed
Released
17 Feb 2026
Status
Generally available
Savings
40%
Speed
no data
Claude Opus 4.7Anthropic $1,233 Cannot confirm 0% 16 Apr 2026 Generally available 0% no data
Release · status · savings · speed
Released
16 Apr 2026
Status
Generally available
Savings
0%
Speed
no data
Claude Opus 4.8Anthropic $1,233 Cannot confirm 0% 28 May 2026 Generally available 0% no data
Release · status · savings · speed
Released
28 May 2026
Status
Generally available
Savings
0%
Speed
no data
Claude Fable 5Anthropic $2,466 Cannot confirm 10% 9 Jun 2026 Generally available -100% no data
Release · status · savings · speed
Released
9 Jun 2026
Status
Generally available
Savings
-100%
Speed
no data
Grok 4xAI no price row covers this date Cannot confirm 0% 9 Jul 2025 Generally available no data
Release · status · savings · speed
Released
9 Jul 2025
Status
Generally available
Savings
Speed
no data
Grok 4 FastxAI $35 Cannot confirm 0% 19 Sep 2025 Generally available 97% no data
Release · status · savings · speed
Released
19 Sep 2025
Status
Generally available
Savings
97%
Speed
no data
Grok 4.1 FastxAI $35 Cannot confirm 0% 19 Nov 2025 Generally available 97% no data
Release · status · savings · speed
Released
19 Nov 2025
Status
Generally available
Savings
97%
Speed
no data
Grok 4.3xAI $185 Cannot confirm 0% Release date not published; listed from 11 Sep 2026, when its price was first read. Generally available 85% no data
Release · status · savings · speed
Released
— (Release date not published; listed from 11 Sep 2026, when its price was first read.)
Status
Generally available
Savings
85%
Speed
no data
Grok 4.5xAI $360 Cannot confirm 0% 16 Jul 2026 Generally available 71% 55 tok/s
Release · status · savings · speed
Released
16 Jul 2026
Status
Generally available
Savings
71%
Speed
55 tok/s
Grok 4.6xAI $380 Cannot confirm 10% 12 Aug 2026 Generally available 69% 55 tok/s
Release · status · savings · speed
Released
12 Aug 2026
Status
Generally available
Savings
69%
Speed
55 tok/s
GPT-5.4OpenAI $685 Cannot confirm 0% 5 Mar 2026 Generally available 44% no data
Release · status · savings · speed
Released
5 Mar 2026
Status
Generally available
Savings
44%
Speed
no data
GPT-5.4 miniOpenAI $205 Cannot confirm 0% 17 Mar 2026 Generally available 83% no data
Release · status · savings · speed
Released
17 Mar 2026
Status
Generally available
Savings
83%
Speed
no data
GPT-5.4 nanoOpenAI $56 Cannot confirm 0% 17 Mar 2026 Generally available 95% no data
Release · status · savings · speed
Released
17 Mar 2026
Status
Generally available
Savings
95%
Speed
no data
GPT-5.5OpenAI $1,370 Cannot confirm 20% 24 Apr 2026 Generally available -11% 85 tok/s
Release · status · savings · speed
Released
24 Apr 2026
Status
Generally available
Savings
-11%
Speed
85 tok/s
GPT-5.5 ProOpenAI $10,890 Cannot confirm 0% 24 Apr 2026 Generally available -783% no data
Release · status · savings · speed
Released
24 Apr 2026
Status
Generally available
Savings
-783%
Speed
no data
GPT-5.6 LunaOpenAI $56 Cannot confirm 0% 9 Jul 2026 Generally available 95% 112 tok/s
Release · status · savings · speed
Released
9 Jul 2026
Status
Generally available
Savings
95%
Speed
112 tok/s
o3-proOpenAI $5,940 Cannot confirm 0% 10 Jun 2025 Deprecated -382% no data
Release · status · savings · speed
Released
10 Jun 2025
Status
Deprecated
Savings
-382%
Speed
no data
o3-miniOpenAI $272 Cannot confirm 0% 31 Jan 2025 Deprecated 78% no data
Release · status · savings · speed
Released
31 Jan 2025
Status
Deprecated
Savings
78%
Speed
no data
o3 deep researchOpenAI no price row covers this date Cannot confirm 0% 26 Jun 2025 Deprecated no data
Release · status · savings · speed
Released
26 Jun 2025
Status
Deprecated
Savings
Speed
no data
DeepSeek-V4-ProDeepSeek $222 Cannot confirm 10% 13 Aug 2026 Generally available 82% 72 tok/s
Release · status · savings · speed
Released
13 Aug 2026
Status
Generally available
Savings
82%
Speed
72 tok/s
Gemini 3.6 FlashGoogle $181 Cannot confirm 0% 21 Jul 2026 Generally available 85% 195 tok/s
Release · status · savings · speed
Released
21 Jul 2026
Status
Generally available
Savings
85%
Speed
195 tok/s
Gemini 3 Flash (Preview)Google no price row covers this date Cannot confirm 55% 17 Dec 2025 Preview no data
Release · status · savings · speed
Released
17 Dec 2025
Status
Preview
Savings
Speed
no data
Gemini 2.5 Flash-LiteGoogle $21 Cannot confirm 0% 22 Jul 2025 Generally available 98% no data
Release · status · savings · speed
Released
22 Jul 2025
Status
Generally available
Savings
98%
Speed
no data
Qwen3.8-MaxAlibaba (Qwen) no price row covers this date Cannot confirm 0% 3 Aug 2026 Generally available no data
Release · status · savings · speed
Released
3 Aug 2026
Status
Generally available
Savings
Speed
no data
Kimi K2.6Moonshot AI (Kimi) $211 Cannot confirm 0% Release date not published; listed from 11 Sep 2026, when its price was first read. Generally available 83% no data
Release · status · savings · speed
Released
— (Release date not published; listed from 11 Sep 2026, when its price was first read.)
Status
Generally available
Savings
83%
Speed
no data
Mistral Large 3Mistral AI $132 Cannot confirm 0% 1 Dec 2025 Generally available 89% no data
Release · status · savings · speed
Released
1 Dec 2025
Status
Generally available
Savings
89%
Speed
no data
MiniMax M2.5MiniMax $62 Cannot confirm 55% Release date not published; listed from 11 Sep 2026, when its price was first read. Deprecated 95% no data
Release · status · savings · speed
Released
— (Release date not published; listed from 11 Sep 2026, when its price was first read.)
Status
Deprecated
Savings
95%
Speed
no data
MiniMax M2.7MiniMax $67 Cannot confirm 0% Release date not published; listed from 11 Sep 2026, when its price was first read. Generally available 95% no data
Release · status · savings · speed
Released
— (Release date not published; listed from 11 Sep 2026, when its price was first read.)
Status
Generally available
Savings
95%
Speed
no data
Claude Mythos 5Anthropic $2,466 Cannot confirm 0% 9 Jun 2026 Preview -100% no data
Release · status · savings · speed
Released
9 Jun 2026
Status
Preview
Savings
-100%
Speed
no data
Claude Mythos 5.1Anthropic $2,395 Cannot confirm 0% 1 Sep 2026 Preview -94% no data
Release · status · savings · speed
Released
1 Sep 2026
Status
Preview
Savings
-94%
Speed
no data
DeepSeek-V4.1-FlashDeepSeek $60 Cannot confirm 0% 10 Sep 2026 Generally available 95% 199 tok/s
Release · status · savings · speed
Released
10 Sep 2026
Status
Generally available
Savings
95%
Speed
199 tok/s
Claude Opus 5Anthropic $1,233 No 85% 24 Jul 2026 Generally available 0% 51 tok/s
Release · status · savings · speed
Released
24 Jul 2026
Status
Generally available
Savings
0%
Speed
51 tok/s
GPT-5.6 SolOpenAI $986 No 85% 9 Jul 2026 Generally available 20% 60 tok/s
Release · status · savings · speed
Released
9 Jul 2026
Status
Generally available
Savings
20%
Speed
60 tok/s
GPT-5.6 TerraOpenAI $559 No 65% 9 Jul 2026 Generally available 55% 87 tok/s
Release · status · savings · speed
Released
9 Jul 2026
Status
Generally available
Savings
55%
Speed
87 tok/s
Gemini 3.7 FlashGoogle $181 No 65% 13 Aug 2026 Generally available 85% 295 tok/s
Release · status · savings · speed
Released
13 Aug 2026
Status
Generally available
Savings
85%
Speed
295 tok/s
Gemini 3.1 Pro (Preview)Google $548 No 85% 19 Feb 2026 Preview 56% 108 tok/s
Release · status · savings · speed
Released
19 Feb 2026
Status
Preview
Savings
56%
Speed
108 tok/s
Claude Sonnet 4.5Anthropic $740 No 55% 29 Sep 2025 Generally available 40% no data
Release · status · savings · speed
Released
29 Sep 2025
Status
Generally available
Savings
40%
Speed
no data
Claude Sonnet 5Anthropic $493 No 65% 30 Jun 2026 Generally available 60% 74 tok/s
Release · status · savings · speed
Released
30 Jun 2026
Status
Generally available
Savings
60%
Speed
74 tok/s
GPT-5OpenAI $425 No 55% 7 Aug 2025 Deprecated 66% no data
Release · status · savings · speed
Released
7 Aug 2025
Status
Deprecated
Savings
66%
Speed
no data
GPT-5 miniOpenAI $85 No 55% 7 Aug 2025 Deprecated 93% no data
Release · status · savings · speed
Released
7 Aug 2025
Status
Deprecated
Savings
93%
Speed
no data
GPT-5 nanoOpenAI $17 No 55% 7 Aug 2025 Deprecated 99% no data
Release · status · savings · speed
Released
7 Aug 2025
Status
Deprecated
Savings
99%
Speed
no data
GPT-5.1OpenAI $425 No 55% 13 Nov 2025 Deprecated 66% no data
Release · status · savings · speed
Released
13 Nov 2025
Status
Deprecated
Savings
66%
Speed
no data
o3OpenAI $446 No 55% 16 Apr 2025 Deprecated 64% no data
Release · status · savings · speed
Released
16 Apr 2025
Status
Deprecated
Savings
64%
Speed
no data
Gemini 2.5 ProGoogle $425 No 55% 17 Jun 2025 Generally available 66% no data
Release · status · savings · speed
Released
17 Jun 2025
Status
Generally available
Savings
66%
Speed
no data
Gemini 2.5 FlashGoogle $105 No 55% 17 Jun 2025 Generally available 91% no data
Release · status · savings · speed
Released
17 Jun 2025
Status
Generally available
Savings
91%
Speed
no data
GLM-4.6Z.ai (GLM) $123 No 55% Release date not published; listed from 11 Sep 2026, when its price was first read. Generally available 90% 37 tok/s
Release · status · savings · speed
Released
— (Release date not published; listed from 11 Sep 2026, when its price was first read.)
Status
Generally available
Savings
90%
Speed
37 tok/s
GLM-5.3Z.ai (GLM) $263 No 65% 14 Aug 2026 Generally available 79% 62 tok/s
Release · status · savings · speed
Released
14 Aug 2026
Status
Generally available
Savings
79%
Speed
62 tok/s
Claude Haiku 4.5Anthropic $247 No 65% 15 Oct 2025 Generally available 80% 78 tok/s
Release · status · savings · speed
Released
15 Oct 2025
Status
Generally available
Savings
80%
Speed
78 tok/s
GPT-5.2OpenAI $595 No 75% 11 Dec 2025 Deprecated 52% 70 tok/s
Release · status · savings · speed
Released
11 Dec 2025
Status
Deprecated
Savings
52%
Speed
70 tok/s
Claude Opus 4.5Anthropic $1,233 No 85% 24 Nov 2025 Generally available 0% 46 tok/s
Release · status · savings · speed
Released
24 Nov 2025
Status
Generally available
Savings
0%
Speed
46 tok/s

How we compare & sources

Method

Every price comes from the vendor's own published pages (pricing pages, docs, launch posts), each with the date we read it. Every capability verdict compares published benchmark results for the same test between a candidate and your baseline model — never a single overall score. Results are graded by how comparable the evidence is (same independent evaluator, same publisher, or separate write-ups), and a capability with no comparable result is always labelled insufficient evidence rather than guessed at.

Price freshness by provider
  • Anthropic: prices checked 1 day ago
  • OpenAI: prices checked 1 day ago
  • Google: prices checked 1 day ago
  • xAI: prices checked 1 day ago
  • DeepSeek: prices checked 0 days ago
  • Moonshot AI (Kimi): prices checked 1 day ago
  • Z.ai (GLM): prices checked 1 day ago
  • Mistral AI: prices checked 1 day ago
  • MiniMax: prices checked 1 day ago
Excluded from this comparison
  • Meta Llama (hosted API): Meta wound down its first-party Llama API on 2026-07-06, so no vendor-published price exists. Third-party hosts publish prices for running the weights, which is the price of a host rather than the price of a model.
  • Qwen3.8-Max (price): The model is listed, but its price is withheld: the official Alibaba Cloud row is labelled "Qwen-Max" and could not be tied to this exact model version.
  • DeepSeek price history before 2026-09-10: Earlier StackCapybara figures were cited to deepseek.ai, which is not a DeepSeek domain. Only prices read from api-docs.deepseek.com are recorded, so DeepSeek has no price history before that date.
  • Off-peak and time-of-day rates: DeepSeek charges half price outside 01:00-04:00 and 06:00-10:00 UTC on weekdays. Estimates use the peak rate, which is the conservative assumption; a workload that runs mostly off-peak costs less than shown.
  • SWE-Bench Pro cost data: The public Scale AI leaderboard (labs.scale.com/leaderboard/swe_bench_pro_public, checked live 2026-09-12) shows resolve rate with confidence interval only. It notes some runs used 'a capped cost limit and turn limit of 50' vs 'uncapped cost... turn limit of 250' and that some rows used the mini-swe-agent harness, but does not publish a structured per-row dollar figure or a consistent agent-version field, so no cost_detail or eval_config could be built without guessing. Scores remain in the dataset unchanged.
  • GPQA Diamond cost data: None of our held GPQA Diamond observations (Anthropic/OpenAI/Epoch AI/Vals AI as evaluators) carry a per-run cost, and GPQA is typically run as a direct model query rather than an agent/harness task, so there is no agent+agent_version to build an eval_config from. Scores remain in the dataset unchanged.
Evidence gaps
  • [anthropic-xai] Exact standard input/output pricing for Grok 4 at its 2025-07-09 launch was not found on an official xAI page still live this session (grok-4 has been dropped from the current docs.x.ai/developers/models pricing table); only third-party aggregator figures ($3/$15) were found.
  • [anthropic-xai] Exact >=200k long-context surcharge $/MTok rate for Claude Sonnet 4.6 (2026-02-17 to 2026-03-13) was not found verbatim on an official page — only Opus 4.6's $10/$37.50 figure was directly quoted from src-opus46-news.
  • [anthropic-xai] Official announcement date and launch pricing history for Grok 4.3 were not found (x.ai/news/grok-4-3 returns 404); current price taken from docs.x.ai/developers/models with effective_from_basis=first_observed.
  • [anthropic-xai] Context window and max output tokens for Claude Opus 4.1, Sonnet 4.5, Opus 4.5, Sonnet 4.6, Opus 4.7, and Opus 4.8 were not stated verbatim on any official page fetched this session (only current-generation models are covered in the live models-overview comparison table), so these fields are null per the no-memory rule despite being widely reported by third parties.
  • [anthropic-xai] Grok 4.3's extended (>=200k) tier pricing is INFERRED by doubling the standard rate per docs.x.ai's stated doubling rule, not verified against an explicit per-model number for grok-4.3 (the page only spelled out the doubled example for grok-4.6).
  • [anthropic-xai] Whether Grok 4, Grok 4 Fast, and Grok 4.1 Fast are still 'active'/GA vs deprecated as of 2026-09-11 was not confirmed — no xAI deprecation/lifecycle page was found (unlike Anthropic's model-deprecations page); Grok 4 no longer appears on the current xAI pricing table.
  • [anthropic-xai] Opus 4.6 system card Table 2.3.A gives one Terminal-Bench 2.0 score for 'GPT-5.2' (64.7%) without stating whether that is Anthropic's Terminus-2 reproduction or a self-reported/OpenAI-harness number; section 2.5's text separately shows Anthropic's own Terminus-2 reproduction of GPT-5.2-Codex was only 57.5%, so the 64.7% table figure likely corresponds to OpenAI's own Codex CLI harness (890 trials) rather than the harmonized Terminus-2 setting used for Opus 4.6 — flagged in comparability_notes but not fully resolved.
  • [anthropic-xai] Grok 4.3 long-context price not published per model (doubling rule only) - withheld.
  • [openai-deepseek] openai.com/index/* launch and pricing posts (GPT-5.6, GPT-5.5, and the 'advancing the price-performance frontier' repricing post) returned HTTP 403 to WebFetch all three times attempted this session. All openai.com-sourced facts here (Sol/Terra/Luna launch date and pre-cut prices, the biology/Terminal-Bench/ExploitBench benchmark figures) are backed only by WebSearch snippets or a forum thread quoting the post, not a direct read of the primary page. Recommend re-attempting with an authenticated fetch method or the Wayback Machine.
  • [openai-deepseek] The GPT-5.6 System Card PDF (https://deploymentsafety.openai.com/gpt-5-6/gpt-5-6.pdf) and the Preview System Card were located but the fetch tool could not parse the binary PDF into text this session (returned raw compressed streams). A proper PDF-extraction pass on this file would likely yield the full official benchmark tables including any OpenAI-reported competitor (Claude/Gemini) comparison numbers requested in scope — this is the single biggest gap versus the ticket's ask for launch-post comparison tables.
  • [openai-deepseek] The exact official-page sentence describing the >272K long-context repricing MECHANISM (whether the higher rate applies to the whole request or only tokens beyond the threshold, and the exact output multiplier) was not found verbatim on developers.openai.com/api/docs/pricing despite two targeted fetch attempts; the 272K threshold itself IS confirmed on that page (via the '<272K context limit' notes column), but third-party aggregators' claim of a 2x input / 1.5x output whole-request multiplier is NOT verified against an official source and was deliberately excluded from the prices array.
  • [openai-deepseek] Reasoning-token billing treatment (are reasoning tokens billed as output tokens at the output rate?) was not explicitly stated in any fetched OpenAI model page text this session; all model records carry reasoning.billing=null.
  • [openai-deepseek] DeepSeek-V3.1 launch pricing numbers were only published as an image on the official news260821 page; could not be read as text this session, so deepseek-v3.1 standard price record has input/output=null.
  • [openai-deepseek] DeepSeek V4 Pro GA (news260813) and V4.1-Flash (news260910) announcement pages likewise embed their pricing tables as images; the numeric values used in the prices array were instead taken from the LIVE api-docs.deepseek.com/quick_start/pricing page (fetched today, 2026-09-11), which reflects current (post all cuts) pricing — treated as effective from the stated announcement dates since DeepSeek's own text says the new pricing 'takes effect' on those dates, but the pre-V4.1-Flash (pre-2026-09-10) DeepSeek-Flash price is not recoverable from any source reached this session.
  • [openai-deepseek] DeepSeek-provided SWE-bench Verified / SWE-bench Pro scores for V4-Pro and V4-Flash (80.6% / 79.0% / 55.4% figures) surfaced only via a third-party inference vendor (fireworks.ai) in WebSearch results, not an official deepseek-ai page or the arxiv technical report (2606.19348) — excluded from observations pending direct confirmation from an official source.
  • [openai-deepseek] gpt-5.3-codex, gpt-5.6-cyber/gpt-5.5-cyber, and realtime/image/video/embedding/fine-tuning model lines were intentionally left out of scope as specialized variants outside the ticket's priority list, despite being present on the official pricing page.
  • [openai-deepseek] Context window, max output, and exact release dates for gpt-5, gpt-5-mini, gpt-5-nano, gpt-5.1, gpt-5.2, gpt-5.4/-mini/-nano were not independently fetched from their individual /api/docs/models/<id> pages this session (only gpt-5.6-sol/terra/luna model pages were fetched) — left null/unverified rather than filled from training-data memory of the pre-GPT-5.6 era, per the no-memory rule.
  • [google-openweight] TOP PRIORITY FINDING: Google's own Gemini 3.8 Flash model card (deepmind.google) and blog.google launch post benchmark it against 'Claude Opus 5' and 'Claude Sonnet 5', NOT 'Claude Opus 4.6'. No official Google page found this session compares Gemini 3.8 Flash to Claude Opus 4.6 specifically. If the site's hypothesis must reference Opus 4.6 exactly, either a different (older) Google comparison needs to be located, or the hypothesis framing needs to be reconciled with what Google is actually benchmarking against as of 2026-09.
  • [google-openweight] Terminal-bench 2.1 score for Gemini 3.8 Flash disagrees between two Google-affiliated readings this session: 89.4% (model card table) vs 90.8% (a WebSearch summary attributed to the blog post). Needs a direct re-fetch of the blog.google page's own benchmark table (not just a search-engine synthesis) to resolve.
  • [google-openweight] The ticket's premise of a contradiction between mistral.ai/pricing and mistral.ai/pricing/api for Mistral Large 3 was NOT reproduced this session -- both pages, fetched independently, showed $0.5/$1.5 per 1M tokens input/output. Either the prior audit caught a since-corrected discrepancy, caught a different model (Medium 3.5?), or the contradiction is intermittent/A-B tested. Worth a follow-up fetch on a different day or via a different region/locale of the pricing pages.
  • [google-openweight] Qwen3.8-Max: open-weight status and exact license are contested between secondary sources (plain Apache-2.0 vs a bespoke revenue-sharing license). Needs a direct Hugging Face page fetch for the actual Qwen3.8-Max (Max-class) weights, not just Qwen3-32B/Qwen3.6-27B, to resolve. Also the qwen3.8-max vs qwen-max API id equivalence on Alibaba Cloud Model Studio was not confirmed with a row explicitly labeled 'qwen3.8-max'.
  • [google-openweight] GLM-4.6 z.ai pricing table column semantics ($0.6 / $0.11 / 'Limited-time Free' / $2.2) were summarized by the fetch tool without the raw table markup being visible to me; the input/cached/output role assignment should be re-verified directly against the live docs.z.ai page.
  • [google-openweight] Gemini 3-pro-preview (original, retired 2026-03-09) launch pricing of $2/$12 (<=200K) and $4/$18 (>200K) was only found via aggregator pages this session, since the model is off Google's live pricing page; not independently confirmed on an archived official Google page (e.g. Wayback Machine snapshot was not checked).
  • [google-openweight] CWE-Bench comparator in the Gemini 3.8 Flash Cyber section of blog.google is anonymized as 'a leading frontier model' (47.8% Pass@1) rather than named -- cannot determine if this is Claude or GPT.
  • [google-openweight] Meta Llama API shutdown (2026-07-06) was confirmed via a WebSearch summary of ai.developer.meta.com/docs/llama-api-deprecation/ after the llama.developer.meta.com URL redirected there; a direct WebFetch of the final ai.developer.meta.com URL returned 404 this session, so the exact page text was not independently re-verified verbatim beyond the search snippet quoted in src-llama-deprecation.
  • [gaps] GPT-6 Astra: exact announcement/GA dateline (Sept 3 vs Sept 4, 2026) is not printed as text on the fetched openai.com/index/gpt-6-astra/ page; only corroborated by secondary press (Axios, CNBC, 9to5Mac).
  • [gaps] GPT-6 Astra and GPT-5.5/5.5 Pro: max output tokens for GPT-5.5 / GPT-5.5 Pro and context window for GPT-5.5 Pro were not stated on the launch post or pricing page fetched this session.
  • [gaps] Mistral Large 3: exact release/announcement date not found on mistral.ai/pricing or mistral.ai/pricing/api (both are pricing-only pages); exact open-weight license name (e.g. Apache-2.0 vs a custom Mistral license) not shown beyond the 'OPEN' category tag.
  • [gaps] Grok 4.3: could not find an entry on the official docs.x.ai/developers/release-notes page fetched this session despite specifically checking; exact release date unconfirmed by an official page (secondary sources disagree: April 17, April 30, or May 6, 2026).
  • [gaps] Grok 4.6: docs.x.ai release notes confirm month (August 2026) and full pricing/threshold breakdown but not the exact day; x.ai/news/grok-4-6 supplies the exact day (August 12, 2026) but a flat headline price without the long-context threshold breakdown — both official pages were used together to reconcile this.
  • [gaps] Kimi K3: exact API-availability date (July 16, 2026) is corroborated only by secondary sources; the official GitHub repo and Hugging Face model card (fetched this session) confirm the license and specs but do not print a launch dateline.
  • [gaps] GLM-5.3: open-weight publish date (secondary sources say August 28, 2026, i.e. 'two weeks after launch') was not independently confirmed on an official page fetched this session — the official blog post only promises weights 'in two weeks'.
  • [gaps] DeepSeek-V4.1-Flash: reasoning/thinking-mode support was not stated on the fetched pricing or changelog pages; unclear whether it is a hybrid-reasoning model like deepseek-v4-pro.
  • [gaps] Gemini 3.8 Flash pricing: the model card's phrasing 'Output: $3.75/1M tokens (with caching)' is ambiguous — output tokens are not typically subject to prompt-cache discounts, so this may actually describe a different pricing mode (e.g. a reduced-effort tier) rather than literal output caching; not reconciled with an explicit Gemini API pricing page this session.
  • [gaps] Claude Opus 5 and Claude Sonnet 5 launch posts: their full benchmark comparison tables (Frontier-Bench v0.1, CursorBench 3.2, ARC-AGI-3, OSWorld 2.0 for Opus 5; the Sonnet 4.6 vs Opus 4.8 comparison chart for Sonnet 5) are rendered as canvas/chart images on anthropic.com and could not be extracted as text or read via screenshot in this session (screenshot capture on the headless browser returned solid black frames).
  • [gaps] GPT-6 Astra benchmark tables: the model's own launch post never includes a Claude Opus 4.6/4.7/4.8 or GPT-5.5 column — competitor columns are limited to GPT-5.6 Sol, Claude Fable 5/5.1, Claude Opus 5, and Gemini 3.8 Flash.
  • [independent] Individual AA sub-benchmark component scores (GPQA Diamond, SciCode, HLE, Terminal-Bench v4.0, AA-LCR v1.1, AA-Briefcase, GDPval-AA v2, AutomationBench-AA, GDP.pdf, CritPt) are known to make up the v4.3 Intelligence Index composite, but their per-model numeric values render as interactive charts on artificialanalysis.ai without an accessible ld+json/table fallback - could not extract exact per-eval scores this session, only the composite Index and AA-Omniscience (which does expose ld+json).
  • [independent] Epoch AI Benchmarking Hub (epoch.ai) was not fetched this session due to time constraints - no GPQA Diamond/FrontierMath/SWE-bench-Verified-on-Epoch's-own-harness data collected.
  • [independent] Scale SEAL leaderboards (scale.com/leaderboard) - SWE-Bench Pro, MCP Atlas, HLE - were not fetched this session due to time constraints.
  • [independent] OSWorld-Verified leaderboard (os-world.github.io) was not fetched this session due to time constraints - no computer-use data collected.
  • [independent] Long-context evals (Context Arena MRCR v2, Fiction.LiveBench) and an independent BrowseComp board were not fetched this session due to time constraints.
  • [independent] METR: could not reliably extract the rendered 50%-time-horizon hour value directly from metr.org/time-horizons/ (it is computed client-side from a logistic fit: coefficient -0.412, intercept 3.912 for 'Claude Opus 4.6 (Inspect)'); the 14.5-hour figure in observations is sourced from a secondary blog (MindStudio) citing METR, not read directly off METR's own chart text/tooltip. Needs primary-source confirmation (e.g. METR's own Opus 4.6 blog post or X thread) if precision matters.
  • [independent] METR has no Gemini 3.8 Flash entry at all as of this session's fetch (metr.org/time-horizons/, last updated 2026-05-08) - cannot compare Gemini 3.8 Flash's autonomous time horizon to Opus 4.6's via METR.
  • [independent] SWE-bench.com's bash-only/Verified leaderboard has no 'Gemini 3.8 Flash' row (only Gemini 3 Flash / Gemini 3 Pro variants) as of this session's fetch - cannot compare directly to Claude 4.6 Opus (75.60% resolved) on this specific board.
  • [independent] Terminal-Bench v4.0 (tbench.ai) has no 'Opus 4.6' row (only Opus 4.7/4.8/5, Fable 5/5.1) as of this session's fetch, despite having Gemini 3.8 Flash (19.1% ± 3.4%) - on-page text search for 'Opus 4.6' returned zero matches.
  • [independent] ARC Prize leaderboard (arcprize.org) has extensive Claude Opus 4.6 data (ARC-AGI-1/2/3, multiple reasoning-effort variants) but no 'Gemini 3.8 Flash' row found via on-page text search as of this session's fetch.
  • [independent] Two rows on arcprize.org are both labeled 'Gemini 3.1 Pro (Preview)' with different data (one has ARC-AGI-1/2 scores + no ARC-AGI-3 data, the other has only an ARC-AGI-3 score) and no visible disambiguation of reasoning effort between them - flagged as ambiguous rather than guessed.
  • [independent] Claude Sonnet 5 (max)'s Artificial Analysis-reported TTFT of 176.14s (vs a 3.67s median for its tier) looks anomalous / possibly a data quality issue on AA's side, but is recorded as-is per the no-correction rule.
  • [independent] Model API ids (e.g. Anthropic's exact versioned model string for Opus 4.6/Opus 5/Sonnet 5) were not confirmed via any provider page this session, so model_id is null throughout; only evaluator-displayed labels are recorded.
  • [independent] METR time horizon for Claude Opus 4.6 withheld: the figure could only be read from a secondary blog.
  • [independent-2] OSWorld-Verified (os-world.github.io, v1) and OSWorld 2.0 (osworld-v2.xlang.ai) leaderboards contain NO Claude Opus 4.6 row at all as of 2026-09-11 — the Anthropic entries jump from Opus 4.5/4.6-era gaps straight to Opus 4.7/4.8/5. Cannot produce any Opus-4.6-vs-candidate computer-use observation from either OSWorld leaderboard version.
  • [independent-2] GPT-6 Astra is absent from: LMArena Text Arena (Overall and Coding, searched via in-page search box, 'No models found'), Vals AI GPQA Diamond, Vals AI MMLU Pro, Context Arena MRCR v2 (8-needle API), and Epoch AI SWE-bench Verified. It IS present on Epoch AI GPQA Diamond, Epoch AI FrontierMath Tiers 1-3 v2, AA Intelligence Index v4.3 components, and Scale SEAL Humanity's Last Exam.
  • [independent-2] Claude Fable 5.1 is absent from Epoch AI's GPQA Diamond leaderboard and from Context Arena's MRCR v2 API despite being present on Epoch FrontierMath, AA, LMArena, Vals, and Scale boards.
  • [independent-2] On Epoch AI's SWE-bench Verified board, only Claude Opus 4.6 (no thinking) and a plain 'DeepSeek v4 (max)' row (build tag unconfirmed) are present among the priority list; Sonnet 5, Gemini 3.8 Flash, Grok 4.6, GPT-6 Astra, Opus 5, Fable 5.1, GPT-5.6 family, Kimi K3, GLM-5.3 are all absent from that specific board as of retrieval — this Epoch board looks sparse/not-yet-backfilled for the newest models.
  • [independent-2] Scale SEAL's SWE-Bench Pro (Public and Private Dataset) boards likewise show almost no priority candidates besides Gemini 3.1 Pro (older) alongside Opus 4.6 — the top rows are last-generation models (Muse Spark 1.1, gpt-5.4, Opus 4.5, Sonnet 4.5); could not confirm whether this reflects the boards not being refreshed since Opus 4.6/5-era releases, or a deliberate freeze. The exact agent harness for the SWE-Bench Pro PRIVATE dataset rows (mini-swe-agent vs. other) was not stated on that specific page (only the PUBLIC dataset page carried the '*Run with mini-swe-agent harness' legend), so harness attribution for the private-set numbers is uncertain.
  • [independent-2] Did not check LMArena WebDev Arena or Vision Arena (goal item 5) due to time budget — only Text Arena Overall and Coding categories were captured with CI/votes.
  • [independent-2] Did not check Vals AI's Finance Agent v2, LegalBench, SWE-bench (archived), Terminal-Bench 2.0, or LiveCodeBench pages for Opus 4.6 pairing — only GPQA Diamond and MMLU Pro were pulled from Vals.
  • [independent-2] Did not check Scale SEAL's VisualToolBench or MultiChallenge boards (named in the goal's evaluator list) due to time budget; only SWE-Bench Pro (Public/Private), MCP Atlas, and Humanity's Last Exam were captured from Scale.
  • [independent-2] Did not check Epoch AI's SciCode or OTIS Mock AIME pages (named in the goal's evaluator list); only GPQA Diamond, SWE-bench Verified, and FrontierMath Tiers 1-3 (v2) were captured from Epoch.
  • [independent-2] AA Intelligence Index v4.3 for the deprecated Claude Opus 4.6 (Adaptive Reasoning, Max Effort) is marked '32*' with an asterisk meaning 'Estimated' — only 4 of the 10 underlying sub-evaluations (HLE, CritPt, AA-Omniscience, AA-LCR v1.1) have actually been re-run on this model; AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, and GDP.pdf show as blank for Opus 4.6 in every comparison table pulled, so the composite figure itself should not be treated as a like-for-like index score against candidates that have all 10 sub-scores populated.
  • [independent-2] Context Arena's MRCR v2 API returns an 'anthropic/claude-opus-4.6' row with reasoning_mode=null (avg_score_overall 0.482, well below the high/medium/low reasoning rows at ~0.72-0.73); this is presumed to be a non-reasoning default call but the API does not label it explicitly as 'non-reasoning', so setting_class was recorded as non_reasoning based on the score pattern rather than an explicit label.
  • [orchestrator] DeepSeek off-peak (time-of-day) discounted rates are not modelled; estimates use the standard rate.
  • [orchestrator] DeepSeek price history before 2026-09-10 is not available from an official DeepSeek page; earlier StackCapybara figures cited deepseek.ai, which is not a DeepSeek domain.
  • ST-392: SWE-bench Multilingual/Lite/Multimodal/Test boards are present in the same saved leaderboard JSON but were not imported this round -- only Verified was in scope for the cost-detail correction. A future pass could extend eval_config/cost_detail to those boards the same way (Multilingual task_count appears to be 300, not 500 -- MiniMax M2.5 shows only 297 completed cost rows there, needs its own reconciliation before import).
  • ST-392: Terminal-Bench's own leaderboard (https://www.tbench.ai/leaderboard, checked live 2026-09-12) publishes a COST column per row (alongside MODEL/AGENT/RESOLUTION RATE/TOKENS), but none of our 77 held Terminal-Bench observations carry cost_per_task_usd. Importing it with the same per-row precision and task-count reconciliation used here for SWE-bench Verified is a dedicated follow-up, not done this round -- flagged via spawn_task.
  • claude-opus-4-1: context_window
  • claude-opus-4-1: max_output not confirmed via an official page fetched this session
  • claude-sonnet-4-5: context_window
  • claude-sonnet-4-5: max_output
  • claude-opus-4-5: context_window
  • claude-opus-4-5: max_output
  • claude-sonnet-4-6: max_output
  • claude-opus-4-7: context_window
  • claude-opus-4-7: max_output
  • claude-opus-4-8: context_window
  • claude-opus-4-8: max_output
  • claude-sonnet-5: Full benchmark table values were only shown as a chart image; only a few numeric HLE/OSWorld-Verified/Firefox Exploit figures for Sonnet 4.6 were printed in text.
  • claude-opus-5: exact SWE-bench Verified % at launch - only seen via aggregator, not a direct quote from src-opus5-news
  • claude-opus-5: Exact numeric Frontier-Bench v0.1 / CursorBench 3.2 / ARC-AGI-3 / OSWorld 2.0 scores were shown only as chart images on the launch post, not extractable as text.
  • grok-4: standard input/output pricing at launch — third-party aggregators claim $3/$15 per MTok but this was not found on an official xAI page fetched this session
  • grok-4: max_output
  • grok-4: current lifecycle status (active/deprecated/retired)
  • grok-4-fast: current lifecycle status
  • grok-4-1-fast: whether a >128k or >200k higher pricing tier exists for this model — not found on an official page this session
  • grok-4-1-fast: current lifecycle status
  • grok-4.3: announced/released_api date
  • grok-4.3: max_output
  • grok-4.3: exact extended (>=200k) tier pricing — inferred, see prices notes
  • grok-4.3: Official release date could not be confirmed from docs.x.ai release notes this session; secondary sources (blogs) give April 17 or April 30 or May 6, 2026, inconsistently.
  • grok-4.5: max_output
  • grok-4.5: current lifecycle status
  • grok-4.6: current lifecycle status
  • gpt-5: context_window
  • gpt-5: max_output
  • gpt-5: reasoning.billing
  • gpt-5-mini: context_window
  • gpt-5-mini: max_output
  • gpt-5-nano: context_window
  • gpt-5-nano: max_output
  • gpt-5.1: context_window
  • gpt-5.1: max_output
  • gpt-5.2: context_window
  • gpt-5.2: max_output
  • gpt-5.4: announced
  • gpt-5.4: released_api
  • gpt-5.4: context_window
  • gpt-5.4: max_output
  • gpt-5.4-mini: announced
  • gpt-5.4-mini: released_api
  • gpt-5.4-mini: context_window
  • gpt-5.4-mini: max_output
  • gpt-5.4-nano: announced
  • gpt-5.4-nano: released_api
  • gpt-5.4-nano: context_window
  • gpt-5.4-nano: max_output
  • gpt-5.4-nano: Exact context window not stated on launch post (pricing page also silent).
  • gpt-5.5: max_output
  • gpt-5.5: reasoning.billing
  • gpt-5.5: long_context_threshold mechanism
  • gpt-5.5: Exact max output token limit not stated on launch post.
  • gpt-5.5-pro: context_window
  • gpt-5.5-pro: max_output
  • gpt-5.5-pro: Context window and max output not stated on launch post; pricing page gives standard-tier price only (no long-context threshold shown for Pro).
  • gpt-5.6-sol: announced (secondary source)
  • gpt-5.6-sol: released_api (secondary source)
  • gpt-5.6-sol: reasoning.billing
  • gpt-5.6-terra: reasoning.billing
  • gpt-5.6-luna: reasoning.billing
  • gpt-6-astra: announced
  • gpt-6-astra: released_api
  • gpt-6-astra: context_window
  • gpt-6-astra: max_output
  • gpt-6-astra: lifecycle status precision
  • gpt-6-astra: Exact announcement day (Sept 3 vs Sept 4) not printed as a dateline in the fetched openai.com page text; taken from secondary reporting (Axios, CNBC).
  • o3: context_window
  • o3: max_output
  • o3-pro: context_window
  • o3-pro: max_output
  • o3-mini: context_window
  • o3-mini: max_output
  • o3-deep-research: pricing (no longer listed)
  • o3-deep-research: context_window
  • o3-deep-research: max_output
  • deepseek-v3.1: license (not independently confirmed for V3.1 specifically)
  • deepseek-v3.1: standard price at launch (image-only on official page)
  • deepseek-v3.1: max_output
  • deepseek-v3.2-exp: license
  • deepseek-v3.2-exp: context_window (carried over from V3.1, not independently re-confirmed)
  • deepseek-v3.2-exp: max_output
  • deepseek-v4-flash: reasoning.billing
  • deepseek-v4-flash: whether 284B/13B params still apply post V4.1 upgrade
  • deepseek-v4-flash: max_output for V4.1 specifically
  • deepseek-v4-pro: max_output
  • deepseek-v4-pro: reasoning.billing
  • deepseek-v4-pro: exact retirement-reversal quote (paraphrased by fetch tool, not re-verified verbatim)
  • gemini-3.8-flash: exact knowledge cutoff date (one summarizer read 'March 2026' from model card, not independently re-verified verbatim)
  • gemini-3.8-flash: whether an >200K-token pricing tier exists for 3.8 Flash (pricing page showed a single flat rate, unlike 3.1 Pro/2.5 Pro)
  • gemini-3.8-flash: The model card's 'with caching' vs 'regular' price labels for both input and output are ambiguous from the extracted text; see prices notes.
  • gemini-3.7-flash: exact context window / max output not independently confirmed on a dedicated 3.7 Flash page; inferred from pricing-page grouping with 3.8/3.6 Flash
  • gemini-3.6-flash: Exact release day within July 2026 not stated on the model card.
  • gemini-3.1-pro-preview: exact max_output token figure not independently re-confirmed on a dedicated model page this session
  • gemini-3-pro-preview: launch pricing not confirmed on an official Google page this session (site no longer serves pricing for a retired preview model); sourced only from aggregator/secondary pages
  • gemini-3-flash-preview: launch pricing not independently re-verified via direct official page fetch
  • gemini-3-flash-preview: context window / max output not confirmed this session
  • gemini-2.5-pro: exact GA date not retrieved this session
  • gemini-2.5-flash-lite: exact GA date not retrieved this session
  • qwen3.8-max: whether Qwen3.8-Max weights are actually published open-weight and under what exact license
  • qwen3.8-max: exact qwen3.8-max vs qwen-max API id equivalence
  • qwen3.8-max: context window not independently confirmed for this specific model id
  • qwen3.8-max: Hosted price withheld: the official Alibaba row is labelled Qwen-Max and could not be tied to this exact model version.
  • kimi-k3: whether kimi-k3 weights are open (K2-series precedent is Modified MIT; K3 not independently confirmed)
  • kimi-k3: release date
  • kimi-k3: Exact API-availability date (as opposed to weight-release date) not printed with a dateline on the official docs/GitHub pages fetched; July 16, 2026 figure is corroborated only by secondary sources for the exact day.
  • kimi-k3: Launch date 2026-07-16 is corroborated only by secondary sources; withheld.
  • glm-4.6: exact column mapping of the z.ai pricing table (input vs cached vs output order) -- values recorded but the $0.6/$0.11/$2.2 triple's role assignment should be re-checked against the live table before use
  • glm-4.6: context window (200K) sourced from a secondary summary of the model, not the official z.ai page itself
  • mistral-large-3: release date (2025-12-01) not confirmed on an official Mistral page this session
  • mistral-large-3: context window 262144 vs a secondary '256k' figure seen elsewhere -- not reconciled
  • mistral-large-3: Exact license name (only the 'OPEN' category tag is shown, not a specific license identifier like Apache-2.0). Context window and max output not listed on the pricing page.
  • mistral-large-3: Release date not found on official pricing pages.
  • minimax-m2.5: context window (204,800) and max output (131,072) figures came from an aggregator, not the official MiniMax docs page fetched this session
  • minimax-m2.7: Modified-MIT commercial-use-restriction claim for M2.7 came only from a secondary summary, not the raw Hugging Face LICENSE file text itself
  • minimax-m2.7: context window / param counts unconfirmed this session
  • deepseek-v4.1-flash: Whether V4.1 Flash supports extended/thinking reasoning mode was not stated on the fetched pricing/changelog pages.
  • glm-5.3: Exact open-weight publish date (secondary sources say 2026-08-28) not confirmed on an official page fetched this session.
Sources (154)

Dataset version 2026-09-12.14, verified through 12 Sep 2026.

How to read this

The three pick cards at the top answer three different questions: the strongest option with enough comparable evidence, the cheapest option confirmed to meet your baseline, and the best value overall. A verdict of "insufficient evidence" is not a bad result: it means no comparable benchmark result exists yet, and the page says so rather than guessing.