LLM API Pricing Reference: Quick Comparison & Best-Use Guide
Evergreen developer API pricing table with SWE-bench scores, best-for notes, and links to StackCapybara model reviews.
Research-based — not hands-on tested. Our scores are editorial judgements compiled from vendor documentation, published pricing and independent user reports. How we review.
This is StackCapybara’s evergreen API pricing reference — expanded with benchmark columns, and per-model profiles. Rates are per 1 million tokens on direct developer APIs.
For long-form routing recipes see LLM API Pricing vs. Quality. For consumer plans see subscription pricing.
Prices in the table below are read from StackCapybara’s shared model-pricing dataset (version 2026-09-12.14 — each vendor’s official pricing page and retrieval date are cited per row) or, where a model is not in that dataset, marked as such rather than hand-typed from memory. Re-check vendor docs before production contracts. Benchmarks cite our flagship research sheet where noted.
Prices checked 26 September 2026 against each vendor’s own pricing page. Claude Opus 5.5 and Claude Fable 5.1 are new since the shared dataset’s last refresh (version 2026-09-12.14), so those two rows below cite Anthropic’s pricing page directly rather than the dataset; every other row still follows the dataset.
API pricing & benchmarks
| Model (API) | In / Out per 1M | SWE Pro | SWE Verified | GPQA | Terminal-Bench |
|---|---|---|---|---|---|
DeepSeek-V4 Flashdeepseek-v4-flash |
See the price cards above Peak/off-peak; this id now routes to DeepSeek-V4.1 Flash |
— | — | — | — |
DeepSeek-V4 Prodeepseek-v4-pro |
See the price cards above Peak/off-peak rate |
— | — | — | — |
Kimi-K2.7-Codekimi-k2.7-code |
Not in the shared dataset Moonshot/Kimi bills in CNY with separate cache-hit / cache-miss input rates, not a flat USD split — see the official pricing link below rather than a converted figure here |
— | — | — | — |
Grok 4.3grok-4.3 |
$1.25 / $2.50 $0.20 / 1M cached input · shared dataset v2026-09-12.14 |
— | — | — | — |
Gemini 3.5 Flashgemini-3.5-flash |
$1.50 / $9.00 $0.15 / 1M cached input (90% off) · shared dataset v2026-09-12.14, effective 2026-05-19 |
55.1% | — | 92.2% | 76.2% |
Claude Sonnet 5 / 4.6claude-sonnet-5 · claude-sonnet-4-6 |
See the price cards above for Claude Sonnet 5claude-sonnet-4-6 (prior gen, still separately priced): $3 / $15 · $0.30 / 1M cached input (90% off) · shared dataset v2026-09-12.14, effective 2026-02-17 |
— | — | 93% | — |
Claude Opus 4.8claude-opus-4-8 |
$5 / $25 $0.50 / 1M cached input (90% off) · shared dataset v2026-09-12.14, effective 2026-05-28 |
69.2% | 88.6% | 93.6% | 74.6% |
Claude Opus 5.5claude-opus-5-5 |
$4 / $20 Anthropic pricing page, retrieved 2026-09-26 — not yet in the shared dataset |
— | — | — | — |
Claude Fable 5.1claude-fable-5-1 |
$10 / $50 Anthropic pricing page, retrieved 2026-09-26 — not yet in the shared dataset |
— | — | — | — |
GPT-5.5 (Standard)gpt-5.5 |
$5 / $30 $0.50 / 1M cached input (90% off) · shared dataset v2026-09-12.14, effective 2026-04-24 |
— | 85.7% | 93.6% | 82.7% |
GPT-5.5 Pro (Reasoning)gpt-5.5-pro |
$30 / $180 No published cached-input rate · shared dataset v2026-09-12.14, effective 2026-04-24 |
— | 85.7% | 93.6% | 82.7% |
Benchmark figures are indicative — drawn from vendor-reported and publicly available results as of 2026. Testing methodologies differ between labs and change over time, so treat these as directional and verify against current published benchmarks before relying on them.
Per-model profiles
Click a model name in the table to jump to its profile. Dedicated single-model reviews are linked when published.
DeepSeek-V4 Flash
The default volume tier when you need millions of cheap API calls. Pair with a router that escalates hard tasks to Pro or a flagship model.
- API ID:
deepseek-v4-flash— retired 2026-09-10; requests are now routed todeepseek-v4.1-flash(DeepSeek-V4.1 Flash), which is what the price cards above reflect. - Pricing: See the current price cards at the top of this page for DeepSeek-V4.1 Flash — billed at a peak and an off-peak rate, not one flat price.
- Context: 128K in / 8K out
- Best for: High-volume extraction, classification, and cheap always-on agent loops under 128K context.
- Benchmarks: No public SWE-bench Pro / GPQA sheet in our verified research set — treat as budget tier.
- Primary review: DeepSeek-V4 API review
- Also mentioned in:
DeepSeek-V4 Pro
Step up from Flash when tool use and syntax quality matter but you are not ready to pay flagship per-token rates.
- API ID:
deepseek-v4-pro - Pricing: See the current price cards at the top of this page for DeepSeek-V4 Pro — billed at a peak and an off-peak rate, not one flat price.
- Context: 128K in / 16K out
- Best for: Stateless coding automation, tool-call chains, and one-shot CI review passes at ~60% lower cost than Sonnet-class APIs.
- Benchmarks: Vendor claims strong MoE coding; SWE-bench Pro not in our verified sheet yet.
- Primary review: DeepSeek-V4 API review (Pro tier)
- Also mentioned in:
Kimi-K2.7-Code
Optimized for agent loops, not single-shot answers. Best when sessions run 10+ tool rounds on the same repo.
- API ID:
kimi-k2.7-code - Pricing: Not in the shared pricing dataset. Moonshot/Kimi’s official pricing page (see “Official pricing sources” below) lists CNY-denominated cache-hit/cache-miss rates for
kimi-k2.7-code, not a flat USD input/output split — check that page directly rather than a converted dollar figure here. Preserved Thinking (no extra replay cost) is a vendor-stated feature, not a price claim. - Context: 256K in / 16K out
- Best for: Long multi-turn terminal/coding sessions where preserved chain-of-thought cuts output-token waste.
- Benchmarks: Moonshot publishes loop-efficiency claims; add SWE-bench row when vendor sheet is verified.
- Primary review: Kimi Code & Kimi-K2.7-Code review
- Also mentioned in:
Grok 4.3
Useful third-rail option when you want API diversity outside OpenAI / Google / Anthropic.
- API ID:
grok-4.3 - Pricing: $1.25 input / $2.50 output per 1M · $0.20 / 1M cached input · shared dataset v2026-09-12.14 (xAI docs, retrieved 2026-09-11)
- Context: 128K in / 8K out
- Best for: Mid-cost general API workloads and X-ecosystem integrations where real-time social context matters.
- Benchmarks: Benchmark sheet not yet verified in StackCapybara research bundle.
- Primary review: Grok 4.3 API review
- Also mentioned in:
Gemini 3.5 Flash
Not a peer flagship — a fast, cheap router tier. Excellent GPQA for the price; SWE Pro mid-tier.
- API ID:
gemini-3.5-flash - Pricing: $1.50 input / $9.00 output per 1M · $0.15 / 1M cached input (90% off) · shared dataset v2026-09-12.14 (Google AI pricing docs, retrieved 2026-09-11), effective 2026-05-19
- Context: 1.04M in / 64K out
- Best for: Speed-tier default: huge context ingestion, multimodal parsing, and search-grounded verification.
- Benchmarks: SWE-bench Verified unverified in our sheet; Pro and GPQA cited in flagship research.
- Primary review: Gemini 3.5 Flash API review
- Also mentioned in:
Claude Sonnet 5 / 4.6
The pragmatic Anthropic tier for interactive coding, refreshed by the June 30, 2026 Claude Sonnet 5 release — near-Opus agentic performance at Sonnet pricing. Terminal workflow review: Claude Code.
- API ID:
claude-sonnet-5·claude-sonnet-4-6 - Pricing: See the current price cards at the top of this page for Claude Sonnet 5 (the cheaper, active-generation rate).
claude-sonnet-4-6is the prior generation, still separately priced at $3 input / $15 output per 1M · $0.30 / 1M cached input (90% off) — shared dataset v2026-09-12.14, effective 2026-02-17. - Context: 1M in / 16K out
- Best for: Fast IDE-style edits, conversational refactors, and daily builder workflows with strong coding accuracy.
- Benchmarks: GPQA estimated from Sonnet-class vendor range; use Opus row for verified SWE Pro ceiling.
- Primary review: Claude Sonnet API review
- Also mentioned in:
Claude Opus 4.8
Record SWE-bench Pro in our verified set. Pay for accuracy when a bad diff is expensive.
- API ID:
claude-opus-4-8 - Pricing: $5 input / $25 output per 1M · $0.50 / 1M cached input (90% off reads) · shared dataset v2026-09-12.14, effective 2026-05-28
- Context: 1M in / 128K out
- Best for: Repo-scale SWE, multi-file architecture refactors, and highest-stakes code reasoning.
- Benchmarks: Verified in StackCapybara flagship research (May 2026).
- Primary review: Claude Opus 4.8 API review
- Also mentioned in:
Need an API key first? Our guide to how to get access to the Claude API covers sign-up, billing and rate limits.
Claude Opus 5.5
Anthropic’s next Opus tier, priced below Claude Opus 5 at launch.
- API ID:
claude-opus-5-5 - Pricing: $4 input / $20 output per 1M · Anthropic pricing page, retrieved 2026-09-26. Not yet in the shared pricing dataset — update this row from the dataset once it lands there.
- Best for: Agentic coding and enterprise work at a lower per-token rate than Claude Opus 5.
- Benchmarks: Not yet in our verified research set.
- Also mentioned in:
Claude Fable 5.1
Anthropic’s most capable widely released model, for the most demanding reasoning and long-horizon agentic work.
- API ID:
claude-fable-5-1 - Pricing: $10 input / $50 output per 1M · Anthropic pricing page, retrieved 2026-09-26. Not yet in the shared pricing dataset — update this row from the dataset once it lands there.
- Best for: Long-running reasoning and agentic tasks where accuracy matters more than per-token cost.
- Benchmarks: Not yet in our verified research set.
- Also mentioned in:
GPT-5.5 (Standard)
Strong Terminal-Bench and GPQA scores. Escalate to Pro only for the hardest shell/OS tasks.
- API ID:
gpt-5.5 - Pricing: $5 input / $30 output per 1M · $0.50 / 1M cached input (90% off) · shared dataset v2026-09-12.14, effective 2026-04-24
- Context: 1M in / 128K out
- Best for: Balanced general API default, math-capable agents, and mixed tool-calling pipelines.
- Benchmarks: SWE Verified mid-range 82.6–88.7% in research; SWE Pro unverified.
- Primary review: GPT-5.5 Standard API review
- Also mentioned in:
GPT-5.5 Pro (Reasoning)
Reserve for migrations, infra repair, and OSWorld-class tasks — not everyday chat.
- API ID:
gpt-5.5-pro - Pricing: $30 input / $180 output per 1M · no published cached-input rate · shared dataset v2026-09-12.14, effective 2026-04-24
- Context: 1M in / 128K out
- Best for: DevOps automation, terminal execution guards, and desktop/OS agents where failure cost dominates token cost.
- Benchmarks: Quality similar to Standard on published suites; pricing is the differentiator.
- Primary review: GPT-5.5 Pro API review
- Also mentioned in:
Review coverage
Dedicated API reviews (linked from the table above):
- DeepSeek-V4 Flash & Pro
- Kimi-K2.7-Code
- Grok 4.3
- Gemini 3.5 Flash
- Claude Sonnet
- Claude Opus 4.8
- GPT-5.5 Standard
- GPT-5.5 Pro
Related guides
- LLM API Pricing vs. Quality — task routing & agent-stack recipes
- GPT-5.5 vs Gemini Flash vs Claude Opus — flagship benchmark deep dive
- Best AI coding agents 2026
- Subscription pricing — Plus / Pro / Advanced plans
Official pricing sources:
DeepSeek ·
Moonshot/Kimi ·
xAI Grok ·
Google Gemini ·
Anthropic ·
OpenAI