LLM API Pricing Snapshot Q4 2026: GPT-6.1 Sol, Opus 5.5, and the Post-Astra Cost Landscape
A comprehensive Q4 2026 pricing reference for every major LLM API: OpenAI, Anthropic, Google, DeepSeek, DashScope, and SiliconFlow. With GPT-6.1 Astra cancelled and GPT-6.1 Sol at $2/$10, the effective cost ceiling has collapsed. We map every tier, cache discount, and batch rate so you can route by budget.
The quick answer: OpenAI's flagship API tier no longer exists. GPT-6.1 Astra was cancelled on September 29, 2026. The $10/$50 price point that OpenAI had held for its top model since GPT-6 Astra is gone. GPT-6.1 Sol now sits at the top of OpenAI's publicly available lineup at $2/$10 per million tokens, the same price as GPT-6 Sol before it. Meanwhile, Claude Opus 5.5 launched at $4/$20, Anthropic's Fable 5.1 holds the extended-reasoning crown at $10/$50, and DeepSeek V4.1 Flash replaced V4 Pro at a fraction of the cost. The pricing floor has dropped, the ceiling has collapsed, and the gap between "frontier" and "mid-tier" has never been narrower.
This page is the full pricing table as of October 3, 2026. Bookmark it, share it with your team, and use it to build your routing config.
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
Sources: OpenAI API Pricing, retrieved 2026-10-03; Anthropic Claude Pricing, retrieved 2026-10-03; DeepSeek Pricing, retrieved 2026-10-03; DashScope Model Pricing, retrieved 2026-10-03; SiliconFlow Pricing, retrieved 2026-10-03; Gemini API Pricing, retrieved 2026-10-03.
The Post-Astra Landscape
Before September 29, 2026, the API pricing map had a clear top. OpenAI charged $10/$50 for GPT-6 Astra. Anthropic charged $10/$50 for Fable 5.1. These were the "if you need the absolute best, pay here" models.
Then OpenAI cancelled GPT-6.1 Astra over safety concerns at DevDay. GPT-6.1 Sol launched the same day at $2/$10, described as delivering "near-Astra intelligence for coding, computer use, and professional work at one-fifth of Astra's standard API prices." That one-fifth comparison is to a model that no longer ships.
The practical result: OpenAI's most capable publicly available model now costs the same as Claude Sonnet 5. The "flagship" price band on the OpenAI side dropped from $10/$50 to $2/$10 overnight. Anthropic still holds $10/$50 with Fable 5.1 and $4/$20 with Opus 5.5, but the competitive pressure from GPT-6.1 Sol at $2/$10 is real.
Master Pricing Table
Prices are per million tokens (input / output) unless noted. Cache and batch columns show the discounted rate where the provider publishes one. A dash means the provider has not published a rate.
Frontier Tier ($4+ input)
| Model | Provider | Input | Output | Cached Input | Batch | Context |
|---|---|---|---|---|---|---|
| Claude Fable 5.1 | Anthropic | $10.00 | $50.00 | $0.25 | $5.00/$25.00 | 1M |
| Claude Fable 5 | Anthropic | $10.00 | $50.00 | $1.00 | $5.00/$25.00 | 1M |
| Claude Opus 5.5 | Anthropic | $4.00 | $20.00 | $0.20 | $2.00/$10.00 | 1M |
Fable 5.1 brought a 75% cut to cache reads ($0.25 vs Fable 5's $1.00), making it significantly cheaper for long-context agentic workloads that hit the cache. With GPT-6.1 Astra cancelled, Fable 5.1 is the only model charging $10/$50 at this tier from a major Western provider.
Mid Tier ($1 -- $4 input)
| Model | Provider | Input | Output | Cached Input | Batch | Context |
|---|---|---|---|---|---|---|
| GPT-6.1 Sol | OpenAI | $2.00 | $10.00 | $0.10 | — | 1M |
| GPT-6 Sol | OpenAI | $2.00 | $10.00 | $0.10 | $1.00/$5.00 | 1M |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | $0.20 | $1.00/$5.00 | 200K |
| Qwen3.8-Max | DashScope | ~$2.00 | ~$6.00 | — | — | 1M |
| GPT-6 Luna | OpenAI | $1.00 | $4.00 | $0.10 | $0.50/$2.00 | 1M |
GPT-6.1 Sol and Claude Sonnet 5 land at identical sticker prices. The key difference: GPT-6.1 Sol offers a 1M context window and $0.10 cached input, while Sonnet 5 caps at 200K context but offers batch pricing at 50% off. Qwen3.8-Max sits in the same band with a notably cheaper output rate (pricing converted from RMB at approximate rates; DashScope also offers promotional discounts that can reduce these further).
Value Tier ($0.10 -- $1 input)
| Model | Provider | Input | Output | Cached Input | Batch | Context |
|---|---|---|---|---|---|---|
| Gemini 3.8 Flash | $0.75 | $3.75 | — | — | 1M | |
| Kimi K3 | Moonshot | ~$0.55 | ~$2.20 | — | — | 256K |
| GPT-6.1 Sol Mini | OpenAI | $0.30 | $1.20 | $0.015 | $0.15/$0.60 | 1M |
| DeepSeek V4.1 Flash | DeepSeek | $0.15 | $0.60 | $0.003 | — | 128K |
| Qwen3.8-Flash | DashScope | ~$0.15 | ~$0.47 | — | — | 256K |
This is where the real action is in Q4 2026. DeepSeek V4.1 Flash replaced V4 Pro on September 14 and brought the cost down to $0.15/$0.60 at off-peak rates (peak rates are 2x). V4 Pro requests now route to V4.1 Flash automatically. Qwen3.8-Flash matches DeepSeek on input pricing through DashScope. Gemini 3.8 Flash sits at $0.75/$3.75 as an introductory rate that Google has said will increase in 2027.
Free and Near-Free Tier
| Model | Provider | Input | Output | Context |
|---|---|---|---|---|
| DeepSeek V4 Flash | DeepSeek | $0.07 | $0.28 | 128K |
| Various open-weight models | SiliconFlow | $0.00 | $0.00 | varies |
SiliconFlow continues to offer free-tier access to open-weight models including Qwen, DeepSeek, and GLM variants. Rate limits apply. DeepSeek V4 Flash remains the cheapest non-free option from a major provider at $0.07/$0.28 off-peak.
Price-Parity Clusters
Three natural clusters emerge where models from different providers land at nearly identical pricing:
The $2/$10 cluster is the most competitive. GPT-6.1 Sol, GPT-6 Sol, and Claude Sonnet 5 all charge $2 input and $10 output. Qwen3.8-Max sits close at ~$2/$6. For a routing layer, these are interchangeable on cost and differentiated only by capability, context window, and regional availability.
The $0.15/$0.60 cluster is the cost-optimization sweet spot. DeepSeek V4.1 Flash and Qwen3.8-Flash both land here. DeepSeek offers peak/off-peak pricing (off-peak is the rate shown; peak is 2x). Qwen3.8-Flash pricing includes DashScope promotional rates that may change.
The $10/$50 cluster is now Anthropic-only. Claude Fable 5.1 and Fable 5 both charge $10/$50. With Astra cancelled, no OpenAI model occupies this band. The question for operators: is the Fable tier worth 5x the cost of GPT-6.1 Sol? For extended-reasoning and agentic workloads that need the absolute best, the answer can be yes, especially with Fable 5.1's $0.25 cache reads.
Cache Economics
Prompt caching has become the single largest cost lever in Q4 2026. The spread between cached and uncached input rates varies wildly:
| Provider | Model | Uncached Input | Cached Input | Savings |
|---|---|---|---|---|
| OpenAI | GPT-6.1 Sol | $2.00 | $0.10 | 95% |
| Anthropic | Opus 5.5 | $4.00 | $0.20 | 95% |
| Anthropic | Fable 5.1 | $10.00 | $0.25 | 97.5% |
| DeepSeek | V4.1 Flash | $0.15 | $0.003 | 98% |
| OpenAI | GPT-6.1 Sol Mini | $0.30 | $0.015 | 95% |
OpenAI's caching is automatic (no API changes required; prefixes that repeat across requests get cached). Anthropic requires explicit cache control headers. DeepSeek caches automatically with a KV cache mechanism. DashScope offers explicit context caching through a separate API.
For workloads where the system prompt, few-shot examples, or document context stays constant across requests, the effective per-token cost drops to 2-5% of the listed input rate. A routing layer that pins stable prefixes and cycles only the user turn can cut costs by 90%+ without changing providers.
Batch Processing Discounts
Not every provider offers batch APIs, and the ones that do have different mechanics:
| Provider | Batch Discount | Turnaround | Notes |
|---|---|---|---|
| OpenAI | 50% off input and output | 24 hours | JSONL upload, async results |
| Anthropic | 50% off input and output | 24 hours | Similar JSONL format |
| DashScope | 50% off | 24 hours | OpenAI-compatible batch endpoint |
| DeepSeek | — | — | No batch API |
| — | — | Batch available via Vertex AI |
If your workload tolerates 24-hour latency, batch processing halves the bill on OpenAI and Anthropic. DashScope matches this discount. DeepSeek has not shipped a batch API.
Decision Tree: Route by Budget
Under $0.50/M input budget → DeepSeek V4.1 Flash ($0.15), Qwen3.8-Flash (~$0.15), DeepSeek V4 Flash ($0.07), or SiliconFlow free tier. Best for high-volume, latency-tolerant workloads. DeepSeek's off-peak pricing makes overnight batch runs extremely cheap.
$0.50 -- $2/M input budget → Gemini 3.8 Flash ($0.75), GPT-6.1 Sol Mini ($0.30), Kimi K3 (~$0.55). Good balance of cost and capability. Gemini 3.8 Flash brings 1M context at this price point.
$2 -- $4/M input budget → GPT-6.1 Sol ($2.00), Claude Sonnet 5 ($2.00), Qwen3.8-Max (~$2.00). The "post-Astra frontier" for most production workloads. Route between these three based on task type: GPT-6.1 Sol for coding and computer use, Sonnet 5 for nuanced writing and analysis, Qwen3.8-Max for Chinese-language tasks.
$4 -- $10/M input budget → Claude Opus 5.5 ($4.00). The frontier reasoning model with the best cost-to-capability ratio. Now the only model in this band after Astra's cancellation.
$10+/M input budget → Claude Fable 5.1 ($10.00). Extended reasoning, long-horizon agentic work, and tasks where cache-read economics ($0.25/M) amortize the high base rate.
TheRouter Configuration
With TheRouter routing OpenAI-compatible requests across these providers, a cost-optimized configuration for Q4 2026 might look like this:
# Route by task complexity and cost target
routes:
- match: { capability: "frontier-reasoning" }
providers:
- model: claude-fable-5.1
weight: 80
- model: claude-opus-5.5
weight: 20
- match: { capability: "general" }
providers:
- model: gpt-6.1-sol
weight: 50
- model: claude-sonnet-5
weight: 30
- model: qwen3.8-max
weight: 20
- match: { capability: "high-volume" }
providers:
- model: deepseek-v4.1-flash
weight: 60
- model: qwen3.8-flash
weight: 40
TheRouter routes OpenAI-compatible requests through configured providers and supports provider/model routing and fallback when live product paths support it.
What to Watch in Q4 2026
OpenAI's next expected release is around October 19. Whether they ship another model at the cancelled Astra's $10/$50 price point or continue to push capability down to the $2/$10 Sol tier will reshape this table.
DeepSeek's planned price increase has been reported but no specific rates or dates have been confirmed. If DeepSeek raises V4.1 Flash above $0.15/$0.60, the cost advantage over Qwen3.8-Flash narrows.
GPT-3.5-turbo-0125 retires October 23. If you are still running this model, the replacement path goes through GPT-6.1 Sol Mini ($0.30/$1.20), which is both cheaper and more capable.
DashScope October 10 sunset wave retires several older Qwen models. See our DashScope October 2026 migration guide for the full list and replacement paths.
FAQ
What is the cheapest frontier-class LLM API in Q4 2026?
GPT-6.1 Sol and Claude Sonnet 5 both cost $2/$10 per million tokens. With Astra cancelled, GPT-6.1 Sol is effectively OpenAI's most capable publicly available model at what used to be "mid-tier" pricing. See our GPT-6.1 Sol integration guide for setup details.
How much does Claude Opus 5.5 cost?
$4 per million input tokens and $20 per million output tokens. Cache reads cost $0.20/M. This is 20% cheaper than Opus 5 ($5/$25) and positions Opus 5.5 as the sole occupant of the $4 tier. See our Claude Opus 5.5 routing guide for migration details.
Why did GPT-6.1 Astra get cancelled?
OpenAI cancelled GPT-6.1 Astra over safety concerns at DevDay on September 29, 2026. The model was expected to cost around $10/$50 and compete with Claude Fable 5.1. Its cancellation leaves Fable 5.1 as the only Western-provider model at the $10/$50 tier. See our GPT-6.1 Sol vs Claude Sonnet 5 comparison for how the mid-tier alternatives stack up.
How does DeepSeek's peak/off-peak pricing work?
DeepSeek charges different rates depending on the time of day (UTC). Off-peak rates are the base rates shown in this article. Peak rates are 2x. For V4.1 Flash, that means $0.30/$1.20 at peak vs $0.15/$0.60 off-peak. If you can schedule batch work during off-peak hours, you cut costs in half. See our DeepSeek V4.1 Flash guide for details.
Which providers support prompt caching?
OpenAI (automatic), Anthropic (explicit cache control), DeepSeek (automatic KV cache), DashScope (explicit context cache API), and Google Gemini (context caching with storage pricing). SiliconFlow does not currently offer prompt caching. See our cross-provider prompt caching comparison for implementation details.
How do I pick between GPT-6.1 Sol, Claude Sonnet 5, and Qwen3.8-Max at similar pricing?
All three sit near $2/M input. GPT-6.1 Sol excels at coding and computer use with a 1M context window. Claude Sonnet 5 is strong on nuanced writing and analysis with 200K context. Qwen3.8-Max is the best option for Chinese-language tasks and offers competitive reasoning at ~$2/$6 through DashScope. Route based on task type, not price. See our mid-tier comparison for benchmarks.
Is SiliconFlow really free?
SiliconFlow offers free-tier access to several open-weight models with rate limits. For production workloads, their paid tiers remove rate limits and add SLA guarantees. The free tier is useful for development, testing, and low-volume applications. See our SiliconFlow free models guide for the current free model list.
Pricing data in this article was gathered on October 3, 2026. LLM API prices change frequently. For the most current rates, check each provider's official pricing page directly. DashScope pricing is converted from RMB at approximate exchange rates. DeepSeek pricing shown is off-peak; peak rates are 2x.