New AI releases: August-September 2026
As of 25 September 2026. This is a dated, sourced log of every AI model release we could verify between 1 August and 25 September 2026: not rumours, not roadmap slides. Eight vendors shipped frontier or near-frontier models in the window, and OpenAI’s GPT-6 Astra landed on 3 September. Every entry below is checked against the vendor’s own announcement. Where we could only find a claim in third-party coverage, we’ve said so.
The short version
- OpenAI’s GPT-6 Astra shipped 3 September. OpenAI describes it as “the most intelligent and aligned model in the world.” It’s OpenAI’s new frontier model, replacing GPT-5.x at the top of the range.
- Anthropic shipped Claude Fable 5.1 and Claude Mythos 5.1 on 1 September. Fable 5.1 holds the same per-token price as Fable 5, but cache reads dropped 75%, cutting typical workload costs by around 25%.
- Anthropic followed with Claude Opus 5.5 on 22 September, at $4/$20 per million tokens (a 20% cut from Opus 5), with performance close to Fable 5.1 on most work.
- Google shipped its third Flash model in six weeks on 2 September: Gemini 3.8 Flash and a security-tuned Gemini 3.8 Flash Cyber.
- xAI shipped two Grok versions in the window (Grok 4.6 on 12 August, Grok 4.7 on 21 September), both at the same $2/$6 per million token pricing.
- DeepSeek and Alibaba both moved on cost. DeepSeek-V4-Pro went GA with peak/off-peak pricing, and Qwen open-sourced weights for a Max-class model for the first time.
- Mistral’s most notable ship was infrastructure, not a headline model: OCR 4.1 went GA, and Medium 3.5 landed in Microsoft Foundry and Copilot Studio.
- Microsoft added both Claude Fable 5.1 and GPT-6 Astra to Copilot Cowork and Copilot Studio, within days of each vendor’s own release.
Every release, dated and sourced
| Date | Vendor | Release | Type | Availability | Source |
|---|---|---|---|---|---|
| 2026-08-03 | Alibaba (Qwen) | Qwen3.8-Max | Flagship LLM (2.4T MoE, 95B active, 1M context) | API (QwenCloud), open weights followed 12 Aug | qwen.ai |
| 2026-08-10 | OpenAI | GPT-5.6-Cyber | Security-specialised model, Daybreak Red tier | Limited to approved defenders via Daybreak Red (vulnerability research, exploit validation, security testing) | community.openai.com |
| 2026-08-12 | xAI | Grok 4.6 | Flagship LLM, 500K context | xAI API, Grok apps | docs.x.ai |
| 2026-08-13 | DeepSeek | DeepSeek-V4-Pro (GA) | Open-weight MoE LLM (1.6T total / 49B active) | App, web, API; peak/off-peak pricing from 16 Aug | api-docs.deepseek.com |
| 2026-08-26 | Alibaba (Qwen) | Qwen3.8-Flash / Qwen3.8-Flash-Next | Efficiency-tier LLM (125B total / 6B active MoE), open weights | QwenCloud API; weights on Hugging Face and ModelScope | huggingface.co |
| 2026-08-31 | Mistral AI | OCR 4.1 (GA) | Document OCR model | API (mistral-ocr-latest) |
docs.mistral.ai |
| 2026-09-01 | Anthropic | Claude Fable 5.1 & Claude Mythos 5.1 | Frontier LLM (Fable 5.1 GA); Mythos 5.1 restricted access | Claude API, AWS, Google Cloud, Microsoft Azure; Mythos via Cyber Verification Program (CVP) and Life Sciences Verification Program (LSVP), US organisations only | anthropic.com |
| 2026-09-02 | Google DeepMind | Gemini 3.8 Flash & Gemini 3.8 Flash Cyber | Mid-tier multimodal LLM + security variant | AI Studio, Gemini API, Gemini Enterprise, Gemini app (Pro/Ultra); Flash Cyber via Fairwind Program | blog.google |
| 2026-09-03 | OpenAI | GPT-6 Astra | Frontier LLM (1.05M context, 128K max output) | Phased rollout: Pro/Enterprise/Business Premium first, then Plus/Business, API and AWS “in the coming days” | community.openai.com |
| 2026-09-04 | Microsoft | GPT-6 Astra added to Copilot Cowork & Copilot Studio | Platform integration, not a new model | Microsoft 365 Copilot, availability varies by region/org | techcommunity.microsoft.com |
| 2026-09-21 | xAI | Grok 4.7 | Flagship LLM, larger base model, same 500K context | xAI API, Grok Build, Cursor and other third-party IDEs | docs.x.ai |
| 2026-09-22 | Anthropic | Claude Opus 5.5 | Frontier LLM, $4/$20 per million tokens | Claude API, AWS, Google Cloud, Microsoft Azure | anthropic.com |
We’re updating this table as new releases land and verify. Check back rather than treating it as a one-off snapshot.
Anthropic: Claude Fable 5.1, Claude Mythos 5.1, and Claude Opus 5.5
Anthropic’s Claude Fable 5.1 (claude-fable-5-1) went generally available on 1 September across the Claude API, AWS, Google Cloud and Microsoft Azure. It’s a same-price refresh of Fable 5, not a new tier: input stays at $10 per million tokens and output at $50, but cache reads dropped from $1.00 to $0.25 per million tokens, a 75% cut. Anthropic’s own figures put the effective saving on typical workloads at around 25%, rising to around 45% on heavy coding and agentic runs where cache reuse is high.
For anyone doing the cost math on a multi-turn agent or a long coding session, that cache-read change matters more than the headline per-token rate. It’s the number that moves when you’re re-reading the same system prompt and tool definitions turn after turn. If you’re comparing this against other frontier options, our LLM API pricing reference tracks the current published rates, and the LLM API cost calculator will run the cache-read math for your own token volumes.
Claude Mythos 5.1 shipped the same day at the same capability tier, but it isn’t a general-access model. Anthropic gates it behind two vetting routes: the Cyber Verification Program (CVP), which covers “certain Opus- and Sonnet-class models with reduced cyber safeguards for defensive security work,” and the Life Sciences Verification Program (LSVP), for professional research and development. Both are currently limited to a set of US organisations. If you’re not enrolled in one of those programs, Fable 5.1 is what you’ll actually be able to call.
Three weeks later, on 22 September, Anthropic released Claude Opus 5.5 (claude-opus-5-5), pricing it at $4 per million input tokens and $20 per million output tokens, a 20% cut from Opus 5, with cache reads down to $0.20 per million tokens. Anthropic’s own material puts its performance close to Fable 5.1 on most work at a lower cost, with output speeds increased by more than 30%. It’s available now on the Claude API, AWS, Google Cloud and Microsoft Azure.
Practically: if you’re already on Fable 5, the upgrade to 5.1 is a drop-in model-ID swap with no repricing risk. If Opus 5 is your daily driver, Opus 5.5 is worth testing purely on the price cut. If you’re evaluating Claude for the first time, our Claude API access guide covers getting a key and the auth basics.
OpenAI: GPT-6 Astra and GPT-5.6-Cyber
GPT-6 Astra is OpenAI’s new frontier model. OpenAI unveiled it on 3 September with a 1.05M-token context window and a 128K max output, pricing at $10/$50 per million input/output tokens ($1 for cached input, $12.50 for cache writes). That’s the same headline rate as Claude Fable 5.1, which puts the two in direct price comparison for the first time. Availability is phased: Pro, Enterprise and Business Premium ChatGPT users got it first, with Plus, Business, the API and AWS following “in the coming days” from launch. OpenAI’s own announcement leans heavily on safety framing, which tracks with reporting that a string of unsanctioned agent-driven incidents in July pushed the company to slow the release down and add more guardrails before shipping.
A month earlier, on 10 August, OpenAI launched GPT-5.6-Cyber, a security-specialised model available only through the Red tier of OpenAI’s Daybreak cyber-defence programme, restricted to approved defenders doing authorised vulnerability research, exploit validation and security testing. If you’re not one of those approved organisations, this one doesn’t change anything for you yet, but it’s a clear sign OpenAI is treating offensive-capable security tooling as an access-gated product line rather than a general model feature.
For teams comparing GPT-6 Astra against the rest of the frontier field on cost and throughput, our GPT-5.6 API review is the closest existing coverage until we publish a dedicated GPT-6 Astra review.
Google DeepMind: the third Flash in six weeks, and a bigger model on the horizon
Google shipped Gemini 3.8 Flash and a security-tuned Gemini 3.8 Flash Cyber on 2 September, the third Flash-tier release in about six weeks following 3.7 Flash and 3.6 Flash’s move to general availability. Google’s own post claims “significant improvements from 3.7 Flash across software engineering, agentic tasks, and critical, multi-step reasoning.” Pricing is introductory: $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026, doubling to $1.50/$7.50 after that. Flash Cyber is aimed specifically at vulnerability detection and automated patching across roughly 20 languages, and, like OpenAI’s cyber model, it’s gated, this time behind Google’s Fairwind Program for trusted defenders.
Separately, on 24 September, DeepMind’s Koray Kavukcuoglu said publicly that Gemini 4 is in the post-training phase and expected before the end of the year. That’s a statement of intent, not a release: we’re not putting it in the table above, but it’s worth knowing a bigger Gemini generation is coming while you’re planning what to build against right now.
If Flash’s price point is what draws you, our Gemini 3.6 Flash API review is still broadly representative of the Flash line’s positioning against Sonnet-tier and GPT-mini-tier competitors.
xAI: two Grok releases in six weeks
Grok 4.6 shipped 12 August and Grok 4.7 followed on 21 September, both holding a 500K token context window (down from Grok 4.3’s 1M, a trade xAI made deliberately for stronger coding and agent performance) and identical $2/$6 per million token pricing below a 200K-token prompt threshold, doubling above it. xAI’s own release notes describe 4.7 as using a larger base model, a longer reinforcement-learning run, and more weight on hard, long-running tasks during training. The model checks its own answers more and works longer on difficult problems than 4.6 did. It’s available in Cursor, Grok Build and the xAI API. Grok 5 remains unshipped: xAI has pushed the target at least five times since mid-2025, and there’s still no confirmed date.
If you’re building against Grok already, the pricing didn’t move between 4.6 and 4.7. This is a quality upgrade at the same cost, the more common shape of release this quarter than a price cut.
DeepSeek and Alibaba (Qwen): the cost story
DeepSeek-V4-Pro reached general availability on 13 August, with new peak/off-peak API pricing landing three days later. Off-peak rates run 50% below peak, which rewards batch and background work scheduled outside high-demand windows. It’s a 1.6T-parameter MoE model with 49B active parameters, open weights, and native support for OpenAI’s Responses API shape, which makes it a comparatively easy swap for teams already building against Codex-style tooling. Our DeepSeek V4 API review covers the preview-era pricing; the GA peak/off-peak structure is the main thing that’s changed since.
Alibaba’s Qwen team had the busier two months. Qwen3.8-Max shipped 3 August as a 2.4-trillion-parameter MoE flagship (95B active, 1M context, native text/image/video input) at roughly $2 input / $6 output per million tokens, and on 12 August, Alibaba open-sourced its weights: the first time a Qwen-Max-class model has shipped as open weights rather than API-only. Then on 26 August, the team released Qwen3.8-Flash as a production API model alongside open weights for its underlying architecture, Qwen3.8-Flash-Next, a 125B-total/6B-active MoE design that Alibaba is explicitly positioning as an early preview of the architecture behind Qwen4 (still in training, no release date announced as of late September). Flash-Next trains at roughly a ninth of the cost of the Qwen3.7-Plus generation it replaces.
The practical read: if licence terms matter to you more than raw benchmark position, Qwen’s two-model spread (Max for capability, Flash-Next for near-free inference) is currently the most aggressive open-weight offer in the field.
Mistral AI: infrastructure over headlines
Mistral’s most concrete ship in the window was OCR 4.1 reaching general availability on 31 August, now served under mistral-ocr-latest. There was no new flagship chat model from Mistral in this window: Mistral Large 3 (December 2025) remains the largest public model in the range. The more relevant development for anyone on Microsoft’s stack is that Mistral Medium 3.5 and OCR 4 were added to Microsoft Foundry, and Medium 3.5 to Copilot Studio, putting Mistral inside the same multi-model picker as Anthropic and OpenAI’s offerings for enterprise Microsoft customers.
Microsoft: a faster route to both Claude and GPT
Microsoft isn’t shipping models, but it’s worth tracking as a distribution channel. Claude Fable 5.1 landed in Copilot Cowork and Copilot Studio on 1 September, the same day Anthropic shipped it, per Microsoft’s own announcement. GPT-6 Astra followed into the same surfaces from 4 September, a day after OpenAI’s own announcement. The model picker in Copilot Cowork now sits inside the prompt box rather than a separate menu, split into GPT and Claude sections, with a reasoning-effort slider to trade speed against depth. If your organisation is on Microsoft 365 and wants access to either Fable 5.1 or GPT-6 Astra without a direct Anthropic or OpenAI contract, Copilot is now a working route to both within days of each vendor’s own release, a meaningfully shorter lag than a year ago.
What we left out
We looked for a Meta Llama 4 open-weights release in this window. Several third-party trackers cited a 12 August ship date, but we couldn’t confirm it on Meta’s own AI blog, which as of this writing still shows the original April 2025 Llama 4 herd post with no August 2026 follow-up. We’re leaving it out rather than repeating an unconfirmed claim; if Meta publishes it, we’ll add it to the table. We also skipped Apple, which had no notable LLM or AI-model release in this window, and a string of coding-tool pricing changes (GitHub Copilot Pro+ pricing, Cursor’s usage-pool restructuring, Windsurf’s rebrand to Devin Desktop back in June) that are real but aren’t model releases, so they don’t belong in a release log.
What we’d actually do
- Building a new agentic coding pipeline today: start with Claude Fable 5.1 or GPT-6 Astra depending on which ecosystem you’re already in. Astra’s context window is larger; Fable 5.1’s cache-read pricing is currently cheaper for high-reuse agent loops. If cost is the deciding factor rather than context, test Claude Opus 5.5 too, since it now undercuts both on price. Compare all three with real workload tokens through the cost calculator before committing.
- Cost-sensitive, high-volume classification or extraction work: Qwen3.8-Flash-Next or DeepSeek-V4-Pro’s off-peak tier. Both are open-weight, both are now genuinely cheap to self-host if you have the GPU budget.
- Security or vulnerability-detection workloads: you’ll need to apply for access. GPT-5.6-Cyber, Gemini 3.8 Flash Cyber and Claude Mythos 5.1 are all gated behind vetting programmes, not open API keys. Budget weeks, not days, for approval.
- Already deep in the Microsoft ecosystem: Copilot Cowork now gets you Claude and GPT frontier models inside the same interface, worth checking before you stand up a separate Anthropic or OpenAI contract.
- Waiting for the next big thing: Gemini 4 is in post-training as of late September with no firm date, and Grok 5 has slipped repeatedly since 2025. Don’t plan a roadmap around either shipping on a specific date.
We haven’t independently benchmarked any of the models above; everything here is vendor-published specs and pricing, not our own test results. Treat the qualitative claims (“most intelligent,” “significant improvements”) as vendor marketing until independent evals catch up.
Last updated 25 September 2026. This table is updated as new releases land and are verified against vendor sources. If you spot something we’ve missed, the sourcing bar above is what we’ll hold it to.