AI Model Context Window Pricing Economics: Cost Per Useful Token Across 1M+ Context Providers
Every frontier model advertises 1M tokens, but filling that window costs anywhere from $0.15 to $10. We break down the real cost per useful token at 128K, 256K, 512K, and 1M context depths across OpenAI, Anthropic, Google, DeepSeek, and DashScope — and show how context-aware routing keeps the bill sane.
A 30-second answer: every frontier model now advertises a 1M-token context window, but the cost of actually using that window varies by more than 70x. Filling 1M tokens of input costs $0.15 on DeepSeek V4.1 Flash (off-peak) and $20.00 on OpenAI GPT-6 Astra (long-context tier). Cache economics, tiered pricing, and output costs compound the spread further. This reference page maps the real cost at four context depths (128K, 256K, 512K, 1M) across the major providers, then shows how context-length-aware routing keeps the bill under control.
This is not a repeat of our context window limits reference, which covers window sizes and effective-context benchmarks. And it is not our prompt caching comparison, which covers caching mechanics. Here we model the dollar cost of filling context at scale — the question most teams skip until the invoice arrives.
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
The 1M context pricing table
All prices are per 1M tokens, USD, as of September 2026. Where a provider uses tiered pricing (different rates above a context threshold), both tiers are shown.
| Provider | Model | Context | Input $/M | Input $/M (long) | Cached input $/M | Output $/M | Output $/M (long) |
|---|---|---|---|---|---|---|---|
| OpenAI | GPT-6 Astra | 1M | $10.00 | $20.00 (>272K) | $1.00 (short) / $2.00 (long) | $50.00 | $75.00 |
| OpenAI | GPT-5.6 Sol | 1M | $4.00 | $8.00 (>272K) | $0.40 / $0.80 | $20.00 | $30.00 |
| OpenAI | GPT-5.6 Luna | 1M | $0.20 | $0.40 (>272K) | $0.02 / $0.04 | $1.20 | $1.80 |
| Anthropic | Claude Fable 5.1 | 1M | $10.00 | flat | $0.25 | $50.00 | flat |
| Anthropic | Claude Opus 4.8 | 1M | $5.00 | flat | $0.50 | $25.00 | flat |
| Anthropic | Claude Sonnet 5 | 1M | $2.00 | flat | $0.20 | $10.00 | flat |
| Gemini 3.8 Flash | 1M | $0.75 | $1.50 (>200K) | $0.075 | $3.75 | $7.50 | |
| DeepSeek | V4.1 Flash | 1M | $0.15 (off) / $0.30 (peak) | flat | $0.003 (off) / $0.006 (peak) | $0.60 (off) / $1.20 (peak) | flat |
| DeepSeek | V4 Pro | 1M | $0.66 (off) / $1.32 (peak) | flat | $0.022 (off) / $0.044 (peak) | $1.98 (off) / $3.96 (peak) | flat |
| DashScope | Qwen3.8-Max | 1M | $2.00 | flat | $0.25 | $6.00 | flat |
Sources: OpenAI pricing (retrieved 2026-09-19), Anthropic pricing (retrieved 2026-09-19), DeepSeek pricing (retrieved 2026-09-19), DashScope pricing (retrieved 2026-09-19), Gemini pricing via third-party aggregators (retrieved 2026-09-19).
Three patterns jump out:
- OpenAI doubles long-context pricing. GPT-5.6 and GPT-6 models charge 2x for input and 1.5x for output above 272K tokens. A 500K-token prompt on GPT-5.6 Sol costs $4.00 for the first 272K and $8.00/M for the rest — blended rate around $5.50/M.
- Anthropic and DeepSeek use flat rates. A 1M-token prompt costs the same per token as a 10K-token prompt. For long-context workloads, this removes the math and the surprise.
- Cache economics dominate at scale. DeepSeek V4.1 Flash cache hits cost $0.003/M off-peak — 98% below the already-cheap cache-miss rate. Anthropic's cache hits on Fable 5.1 cost $0.25/M, a 97.5% reduction from the $10.00/M base.
What does it actually cost to fill a context window?
The headline pricing table shows per-million-token rates. Here is what you pay when you fill the window to a specific depth — input cost only, no caching, no output.
| Context depth | DeepSeek V4.1 Flash (off-peak) | GPT-5.6 Luna | Gemini 3.8 Flash | Qwen3.8-Max | Claude Sonnet 5 | GPT-5.6 Sol | Claude Opus 4.8 | GPT-6 Astra |
|---|---|---|---|---|---|---|---|---|
| 128K tokens | $0.02 | $0.03 | $0.10 | $0.26 | $0.26 | $0.51 | $0.64 | $1.28 |
| 256K tokens | $0.04 | $0.05 | $0.19 | $0.51 | $0.51 | $1.02 | $1.28 | $2.56 |
| 512K tokens | $0.08 | $0.10 | $0.58 | $1.02 | $1.02 | $2.88 | $2.56 | $7.68 |
| 1M tokens | $0.15 | $0.20 | $1.13 | $2.00 | $2.00 | $5.60 | $5.00 | $14.60 |
Notes: GPT-5.6 Sol and GPT-6 Astra 1M figures use blended rates (first 272K at short-context rate, remaining 728K at long-context rate). Gemini 3.8 Flash 512K and 1M use blended rates (first 200K at $0.75, rest at $1.50). DeepSeek off-peak rates used throughout.
The 71x spread between DeepSeek V4.1 Flash ($0.15) and GPT-6 Astra ($14.60) at 1M tokens is not an artifact of comparing flagship to budget. Both models have 1M-token windows. The gap reflects architecture, market positioning, and tiered pricing policy.
For a workload that sends 100 requests per day at 500K tokens of context each, the monthly input-only bill ranges from roughly $240 (DeepSeek V4.1 Flash off-peak) to $23,000 (GPT-6 Astra). That is a team-hire-level difference.
Cache economics change everything at scale
Raw input pricing is only half the story. If your workload has repeatable prefixes — system prompts, tool schemas, document chunks used across multiple requests — caching collapses the effective input cost.
| Provider | Model | Cache miss $/M | Cache hit $/M | Savings on hit | Break-even: hits to recover one write |
|---|---|---|---|---|---|
| DeepSeek | V4.1 Flash (off-peak) | $0.15 | $0.003 | 98% | 1 (no write premium) |
| Anthropic | Claude Sonnet 5 | $2.00 | $0.20 | 90% | ~1.2 (5m write at $2.50) |
| Anthropic | Claude Fable 5.1 | $10.00 | $0.25 | 97.5% | ~1.3 (5m write at $12.50) |
| OpenAI | GPT-5.6 Sol | $4.00 | $0.40 | 90% | ~1.4 (write at $5.00) |
| Gemini 3.8 Flash | $0.75 | $0.075 | 90% | context caching free through Dec 2026 | |
| DashScope | Qwen3.8-Max | $2.00 | $0.25 | 87.5% | varies by cache mode |
The pattern: if you hit cache more than once or twice per write, every provider's effective input cost drops by 87–98%. DeepSeek's implicit caching with no write premium makes it the simplest to adopt — you do not need to structure your prompts for caching, it happens automatically.
For a document-review workflow that processes 50 contracts against the same 200K-token system prompt and schema, caching turns a $20.00/day Claude Opus 4.8 bill into a $2.50/day bill (one cache write + 49 cache hits). On DeepSeek V4.1 Flash, the same workload drops from $1.50/day to $0.16/day.
The cost-per-useful-token problem
Not all context is useful context. A 1M-token prompt where 60% is padding, boilerplate, or context the model ignores effectively costs 2.5x the headline rate per useful token.
Research from BenchLM (retrieved 2026-09-19) puts it plainly: advertised context and effective context are different numbers. The gap between them grows as context length increases. At 128K tokens, most frontier models maintain strong recall. At 500K+, recall degrades for some models, which means tokens at the far end of the window may be paid for but not used.
This creates a second pricing lever beyond the per-token rate: context utilization quality. A model that costs 3x more per token but maintains 95% recall at 500K may deliver a lower cost per useful token than a model at one-third the price with 60% recall at the same depth.
We do not have standardized cross-provider effective-context benchmarks at every depth as of September 2026. What we do have:
- DeepSeek V4 Pro reports strong NIAH (needle-in-a-haystack) scores at 1M
- Claude Opus 4.8 and Fable 5 maintain recall across the full 1M window in Anthropic's published evaluations
- Google Gemini 3.1 Pro showed the strongest effective-context scores at 500K–1M in BenchLM's April 2026 evaluation
- OpenAI GPT-5.5 scored highest on LongBench v2 and MRCRv2 (128–256K) in the same evaluation
The practical takeaway: do not pick the cheapest per-token model for long-context work without testing recall at your actual context depth. A $0.15/M model with 50% recall at 500K costs more per useful token than a $2.00/M model with 95% recall.
Output costs: the overlooked multiplier
Input pricing gets all the attention. But for workloads that generate long outputs — report drafting, code generation, document translation — output pricing matters more.
| Model | Output $/M (short) | Output $/M (long) | Max output tokens |
|---|---|---|---|
| DeepSeek V4.1 Flash | $0.60 (off-peak) | flat | 384K |
| GPT-5.6 Luna | $1.20 | $1.80 (>272K) | 128K |
| Gemini 3.8 Flash | $3.75 | $7.50 (>200K) | 64K |
| Claude Sonnet 5 | $10.00 | flat | 64K |
| Qwen3.8-Max | $6.00 | flat | 32K |
| GPT-5.6 Sol | $20.00 | $30.00 (>272K) | 128K |
| Claude Opus 4.8 | $25.00 | flat | 64K |
| GPT-6 Astra | $50.00 | $75.00 (>272K) | 128K |
DeepSeek V4 Pro's 384K output ceiling is structurally unique — 3–6x larger than peers. For workflows that need a full document draft in a single call, this is not a price advantage, it is a capability advantage that no amount of routing fixes for other providers.
A workload generating 50K tokens of output per request at 100 requests/day: DeepSeek V4.1 Flash costs $90/month (off-peak output). Claude Sonnet 5 costs $1,500/month. GPT-6 Astra costs $7,500/month. The output cost alone can exceed the input cost for generation-heavy workloads.
Context-aware routing strategy
Given the pricing spread, the most effective cost control is context-length-based routing: send requests to different providers based on the input context depth.
A simple rule set:
| Context depth | Route to | Rationale |
|---|---|---|
| <32K tokens | Any provider (quality-first) | Cost differences are negligible at low context |
| 32K–128K tokens | Provider with best quality/cost for your use case | Moderate cost; provider quality differences start to matter |
| 128K–256K tokens | Prefer flat-rate providers (Anthropic, DeepSeek) or cache-heavy workflows | Tiered providers (OpenAI, Google) start charging premiums |
| 256K–1M tokens | DeepSeek or Anthropic with caching | Tiered providers charge 2x; flat-rate providers maintain the same rate |
A gateway like TheRouter can implement this routing at the infrastructure level. The application sends every request to the same OpenAI-compatible endpoint; the gateway routes based on the token count in the request. Combined with fallback routing, this means long-context requests automatically land on the cheapest adequate provider, with fallback to a more expensive provider if the primary is unavailable.
We described this pattern in more detail in our cost optimization routing guide. The context-length dimension adds a routing signal that most cost optimization guides miss.
Real workload examples
RAG with 200K retrieval context
A retrieval-augmented generation pipeline passes 200K tokens of retrieved chunks plus a 2K system prompt per query. 500 queries/day.
- DeepSeek V4.1 Flash (off-peak, cached system prompt): $0.03/query input + $0.003 cached prefix. ~$465/month.
- Claude Sonnet 5 (cached system prompt): $0.40/query input + $0.004 cached prefix. ~$6,060/month.
- GPT-5.6 Sol: $0.81/query input (all short-context tier). ~$12,150/month.
Code repository analysis at 500K tokens
A code-review agent loads a 500K-token repo snapshot per review. 20 reviews/day.
- DeepSeek V4.1 Flash (off-peak): $0.08/review. ~$48/month.
- Claude Opus 4.8: $2.56/review. ~$1,536/month.
- GPT-6 Astra (blended): $7.68/review. ~$4,608/month.
Document processing at 1M tokens
A legal document processor ingests full contracts at 800K–1M tokens. 10 documents/day.
- DeepSeek V4.1 Flash (off-peak): $0.15/doc. ~$45/month.
- Claude Sonnet 5: $2.00/doc. ~$600/month.
- GPT-5.6 Sol (blended): $5.60/doc. ~$1,680/month.
In every scenario, context-aware routing to DeepSeek for long-context work and a higher-quality model for short-context reasoning produces the best cost/quality ratio.
FAQ
How much does it cost to fill a 1M-token context window?
It ranges from $0.15 (DeepSeek V4.1 Flash off-peak) to $20.00 (OpenAI GPT-6 Astra long-context tier) per 1M input tokens. The 130x+ spread reflects model tier, architecture, and pricing policy, not just quality differences.
Do all providers charge the same rate regardless of context length?
No. OpenAI doubles input and increases output prices for GPT-5.6 and GPT-6 models above 272K tokens. Google Gemini doubles rates above 200K tokens. Anthropic and DeepSeek use flat rates across the full 1M window. DashScope uses tiered pricing for some Qwen models above 128K tokens.
What is cost per useful token?
Cost per useful token accounts for context utilization quality. If a model ignores 40% of the tokens in a long context window, your effective cost per useful token is roughly 1.67x the raw per-token price. A cheaper model with poor long-context recall can cost more per useful token than an expensive model with strong recall.
How does prompt caching reduce long-context costs?
Caching stores repeated prompt prefixes so subsequent calls pay the cache-hit rate instead of full input pricing. Anthropic charges 10% of base input on cache hits (2.5% on Fable 5.1). OpenAI charges 10%. DeepSeek charges as low as 2% (off-peak cache hit). For workloads with stable system prompts or shared document prefixes, caching reduces effective input cost by 87–98%.
Can routing reduce context window costs?
Yes. Context-aware routing sends short-context requests to any provider and routes long-context workloads to the cheapest provider with adequate quality. A gateway like TheRouter can route based on input token count without application code changes. Combined with model fallback, this keeps long-context costs low while maintaining availability.
Should I always use the cheapest long-context provider?
Not necessarily. Cost per useful token matters more than cost per raw token. If the cheapest provider has poor recall at your context depth, you may need more retries or get lower-quality output — both of which increase total cost. Test recall at your actual context depth before committing to a provider for long-context work.
Pricing data retrieved September 19, 2026. Prices change; verify against provider pricing pages before making procurement decisions.