September 2026 AI Model Launches: Complete API Routing Changelog for GPT-6, Fable, Gemini 3.8, and V4.1
Five major text-generation models shipped in the first two weeks of September 2026. This reference page covers every launch, summarizes pricing and specs, and links to the dedicated routing guide for each model.
September 2026 packed five major text-generation model launches into its first ten days. If you operate an API routing layer, every one of these models changes the math on at least one routing decision. This page is a single reference for all five launches, with specs, pricing, and a link to the dedicated guide for each.
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
September 2026 Launch Timeline
| Date | Model | Provider | Input $/M | Output $/M | Context | Highlights |
|---|---|---|---|---|---|---|
| Sep 1 | Claude Fable 5.1 | Anthropic | $10.00 | $50.00 | 1M | 75% cache-read discount ($2.50/M cached input) |
| Sep 1 | Claude Mythos 5.1 | Anthropic | — | — | — | Research-only; not yet API-available |
| Sep 2 | Gemini 3.8 Flash | $0.75 | $3.75 | 1M | Promo pricing through Dec 31 2026 | |
| Sep 2 | Qwen3.8-Max-0902 | Alibaba / DashScope | ~$1.69 | ~$5.07 | 1M | Snapshot update; improved coding benchmarks |
| Sep 4 | GPT-6 Astra | OpenAI | $10.00 | $50.00 | 256K | 128K max output; cached input $1.00/M |
| Sep 10 | DeepSeek V4.1 Flash | DeepSeek | $0.15 | $1.10 | 128K | Off-peak pricing; peak is $0.30/$2.20 |
Prices listed are standard API rates where available. Promotional, batch, and off-peak discounts are noted in each model's section below.
GPT-6 Astra
GPT-6 Astra is OpenAI's September 2026 frontier model. It launched on September 4 with a 256K context window, 128K max output tokens, and pricing at $10/$50 per million tokens.
Key specs
- Model ID:
gpt-6-astra - Context window: 256,000 tokens
- Max output: 128,000 tokens
- Pricing: $10/M input, $50/M output, $1/M cached input
- Batch pricing: $5/M input, $25/M output (50% discount)
- Effort levels: Low, Medium, High (reasoning compute scaling)
Routing implications. GPT-6 Astra is 13x more expensive on output than Gemini 3.8 Flash. Route complex reasoning, multi-step agentic workflows, and tasks requiring the largest output window here. For commodity classification, extraction, or summarization, the flash tier makes more sense.
What changed from GPT-5.6 Sol. GPT-6 Astra introduces effort-level controls and a larger output window. If you relied on GPT-5.6 Sol's programmatic tool calling, that feature carries forward. Pricing increased from $7.50/$30 to $10/$50.
For a full integration walkthrough, see our GPT-6 Astra API routing guide.
Claude Fable 5.1
Anthropic shipped Claude Fable 5.1 on September 1, alongside the research-focused Mythos 5.1. Fable 5.1 keeps the same headline pricing as Fable 5 ($10/$50) but introduces a significant cache-read discount.
Key specs
- Model ID:
claude-fable-5-1-20260901 - Context window: 1,000,000 tokens
- Max output: 128,000 tokens (with extended thinking)
- Pricing: $10/M input, $50/M output
- Cache read: $2.50/M (75% discount from standard input)
- Cache write: $12.50/M
- Batch pricing: $5/M input, $25/M output (50% discount)
Routing implications. The 75% cache-read discount is the standout feature. If your workload involves repeated system prompts, RAG contexts, or multi-turn conversations with stable prefixes, effective input cost drops from $10/M to $2.50/M. That makes Fable 5.1 cheaper than GPT-6 Astra on cached input while matching on output pricing.
What changed from Fable 5. Benchmark improvements across coding and agentic tasks. The cache-read discount is new. Extended thinking output limits increased. Model ID changed from claude-fable-5-20260601 to claude-fable-5-1-20260901.
For cache economics and setup details, see our Claude Fable 5.1 routing guide.
Gemini 3.8 Flash
Google released Gemini 3.8 Flash on September 2, three weeks after Gemini 3.7 Flash. It ships at the same promotional pricing and targets cost-effective production workloads.
Key specs
- Model ID:
gemini-3.8-flash - Context window: 1,000,000 tokens
- Max output: 65,536 tokens
- Pricing: $0.75/M input, $3.75/M output (promotional through Dec 31 2026)
- Post-promo pricing: $0.15/M input (128K), $0.075/M (over 128K) — see Google pricing page
- Thinking mode: Supported
Routing implications. At $0.75/$3.75, Gemini 3.8 Flash sits in the flash cost tier alongside DeepSeek V4.1 Flash. It is 13x cheaper on output than GPT-6 Astra. Route high-volume tasks here when latency and cost matter more than peak reasoning depth.
What changed from Gemini 3.7 Flash. Improved benchmark scores across coding and STEM tasks. The pricing stays the same. Context window unchanged at 1M tokens. Computer-use capabilities added (see our Gemini 3.5 Flash computer use guide for the pattern).
For a full routing guide, see our Gemini 3.8 Flash routing guide.
DeepSeek V4.1 Flash
DeepSeek released V4.1 Flash on September 10, a distilled variant of DeepSeek V4.1 that targets speed and cost.
Key specs
- Model ID:
deepseek-chat(auto-routes to V4.1 Flash when available) - Context window: 128,000 tokens
- Max output: 16,384 tokens
- Off-peak pricing: $0.15/M input, $1.10/M output
- Peak pricing: $0.30/M input, $2.20/M output
- Cache hit: $0.003/M input (off-peak), $0.007/M (peak)
- Cache miss: $0.15/M input (off-peak)
Routing implications. DeepSeek V4.1 Flash is the cheapest option in this changelog at $0.15/$1.10 off-peak. The peak/off-peak pricing model adds a scheduling dimension. If your workload can tolerate off-peak windows, this is the lowest-cost path for commodity tasks. The cache-hit pricing at $0.003/M is effectively free.
What changed from V4 Flash. Benchmark improvements across coding tasks. Price reductions on both input and output. The distillation from V4.1 brought better speed with comparable quality to V4 Flash.
For integration details, see our DeepSeek V4.1 Flash complete guide.
Qwen3.8-Max-0902
Alibaba released a new snapshot of Qwen3.8-Max on September 2, the first point update to its 2.4-trillion-parameter flagship since its August 3 launch.
Key specs
- Model ID:
qwen3.8-max-2026-09-02(also available asqwen3.8-max) - Context window: 1,000,000 tokens
- Max output: 16,384 tokens (32,768 with thinking mode)
- Pricing: ~¥12/M input,
¥36/M output ($1.69/$5.07 at current rates) - Thinking mode: Supported via
enable_thinking: true
Routing implications. Qwen3.8-Max-0902 is the most cost-effective frontier-class model for operators who need DashScope coverage. Pricing sits between the flash tier and the Western frontier tier. If you route through DashScope's OpenAI-compatible endpoint, no code changes are needed — the snapshot update is automatic if you use the unversioned qwen3.8-max model ID.
What changed from the August 3 launch. Improved coding performance. SWE-Bench scores reportedly higher. Same pricing, same API surface.
For a complete guide, see our Qwen3.8-Max API guide.
Cross-Model Comparison
| GPT-6 Astra | Fable 5.1 | Gemini 3.8 Flash | V4.1 Flash | Qwen3.8-Max-0902 | |
|---|---|---|---|---|---|
| Input $/M | $10.00 | $10.00 | $0.75 | $0.15 | ~$1.69 |
| Output $/M | $50.00 | $50.00 | $3.75 | $1.10 | ~$5.07 |
| Cache read | $1.00/M | $2.50/M | $0.19/M | $0.003/M | — |
| Context | 256K | 1M | 1M | 128K | 1M |
| Max output | 128K | 128K | 65K | 16K | 16K |
| Batch discount | 50% | 50% | — | — | — |
| OpenAI-compat | Native | Yes | Yes | Yes | Yes (DashScope) |
Reading the table. The frontier tier (GPT-6 Astra, Fable 5.1) charges 13x more on output than the flash tier (Gemini 3.8 Flash, V4.1 Flash). Qwen3.8-Max sits in between. Your routing config should reflect this gap: send complex reasoning to the frontier tier, and high-volume commodity tasks to the flash tier.
Routing Decision Tree
- Does the task need deep multi-step reasoning? Route to GPT-6 Astra or Claude Fable 5.1. If the workload has repeated context, prefer Fable 5.1 for its cache-read discount.
- Is the task high-volume extraction, classification, or summarization? Route to Gemini 3.8 Flash or DeepSeek V4.1 Flash. V4.1 Flash is cheaper but has peak/off-peak pricing and a smaller context window.
- Do you need DashScope or a Chinese provider? Route to Qwen3.8-Max via DashScope's OpenAI-compatible endpoint.
- Need a fallback chain? A practical setup: primary
gpt-6-astra→ fallbackclaude-fable-5-1-20260901→ last-resortgemini-3.8-flash. This gives you frontier reasoning with flash-tier resilience.
For detailed fallback configuration, see our LLM API fallback routing guide.
What We Did Not Cover
- Claude Mythos 5.1 launched alongside Fable 5.1 but remains research-access only with no public API pricing. We will cover it when API access opens.
- Provider-specific benchmarks. Each vendor reports scores on different evaluation suites. We did not attempt apples-to-apples benchmark normalization in this changelog. See the individual guides linked above for benchmark details.
- TheRouter live model availability. Model IDs mentioned here reflect upstream provider availability. Check your routing gateway's model list for current coverage.
FAQ
Which September 2026 model should I use for coding agents?
GPT-6 Astra and Claude Fable 5.1 both score well on coding benchmarks. If your agent uses long system prompts that repeat across calls, Fable 5.1's cache-read discount makes it cheaper over time. Qwen3.8-Max-0902 is a strong option if you need DashScope coverage. See our coding agent routing comparison for a broader view.
Is DeepSeek V4.1 Flash production-ready?
V4.1 Flash launched as a public release on September 10. It is OpenAI-compatible and available via DeepSeek's API. The peak/off-peak pricing model and 16K max output are the main constraints for production use. See our DeepSeek V4.1 Flash guide for setup details.
How does Gemini 3.8 Flash compare to DeepSeek V4.1 Flash?
Both sit in the flash cost tier, but Gemini 3.8 Flash has a 1M context window versus V4.1 Flash's 128K. Gemini is $0.75/$3.75 (promotional) while V4.1 Flash is $0.15/$1.10 (off-peak). V4.1 Flash is cheaper, Gemini has a bigger context. See our GPT-6 Astra vs Gemini 3.8 Flash comparison for routing details.
When will promotional pricing end for Gemini 3.8 Flash?
Google's promotional pricing of $0.75/$3.75 per million tokens runs through December 31, 2026. After that, pricing reverts to the standard rate. Check the Google pricing page for updates.
Can I call all five models from a single OpenAI SDK client?
Yes. All five models support OpenAI-compatible request formats. You set the base_url to point at each provider's endpoint (or your routing gateway) and change the model parameter. See our OpenAI-compatible API providers reference for endpoint details across all providers.