Claude Fable 5.1 Prompt Caching Economics: Breakeven Math, Request Patterns, and Routing Strategies
Fable 5.1 cache reads dropped 75% to $0.25/MTok — but only the cache-read line moved. We break down the breakeven math, show which request patterns actually save money, and explain how cache-aware routing decisions interact with fallback to cheaper models.
Claude Fable 5.1 Prompt Caching Economics: Breakeven Math, Request Patterns, and Routing Strategies
Fable 5.1 cache reads cost $0.25 per million tokens — 75% less than Fable 5's $1.00. That single line-item change turns cache-heavy workloads roughly 25% cheaper on average, and up to 45% cheaper for long-running agents. But the headline price ($10 input, $50 output) did not move. If your application never caches, your Fable 5.1 bill is identical to Fable 5, to the cent.
We wrote this guide because cache economics now determine whether Fable 5.1 or a cheaper model is the right routing target for a given workload. The answer depends on your cache-hit ratio, your context length, and how many turns your agent takes — not on the 75% number alone.
Sources: Anthropic pricing page, retrieved 2026-09-14; Anthropic prompt caching docs, retrieved 2026-09-14; Anthropic Fable 5.1 announcement, retrieved 2026-09-14; VentureBeat coverage, retrieved 2026-09-14.
TL;DR: what moved and what did not
| Line item | Fable 5 | Fable 5.1 | Change |
|---|---|---|---|
| Base input | $10 / MTok | $10 / MTok | None |
| 5-min cache write | $12.50 / MTok | $12.50 / MTok | None |
| 1-hour cache write | $20 / MTok | $20 / MTok | None |
| Cache read | $1.00 / MTok | $0.25 / MTok | -75% |
| Output | $50 / MTok | $50 / MTok | None |
The 0.025x multiplier on cache reads is unique to Fable 5.1 and Mythos 5.1. All other Claude models use a 0.1x multiplier. That asymmetry is the entire economics story.
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
How Anthropic prompt caching works
Anthropic caches the key/value tensors from the attention layers for a contiguous prefix of your prompt. On subsequent requests that share the same prefix, the model skips recomputation of those tokens and charges the cache-read rate instead of the base input rate.
Two modes are available:
- Automatic caching: Add
cache_control: {"type": "ephemeral"}at the request level. The system places the cache breakpoint at the end of the last cacheable block and advances it as conversations grow. - Explicit breakpoints: Place
cache_controlon individual content blocks for precise control over what gets cached.
Key constraints:
- Minimum cacheable prefix: 1,024 tokens for Haiku models, 2,048 tokens for all others (including Fable 5.1).
- TTL options: 5-minute ephemeral (1.25x input price to write) or 1-hour extended (2x input price to write). Cache reads refresh the TTL automatically.
- Prefix-only: Caching is strictly prefix-based. If variable content appears before your static context, the cache will not trigger.
The usage response includes cache_creation_input_tokens and cache_read_input_tokens so you can track exactly how many tokens hit the cache versus how many were written fresh.
Cache pricing math: the five token buckets
Every Fable 5.1 request produces up to five billable token categories. Understanding them is necessary before you can calculate breakeven:
- Uncached input tokens — billed at $10/MTok. These are tokens that appear after your cached prefix, or on a first request before the cache is populated.
- 5-minute cache write tokens — billed at $12.50/MTok (1.25x input). You pay this on the first request that establishes the cached prefix, or when the cache has expired.
- 1-hour cache write tokens — billed at $20/MTok (2x input). Same as above but with a longer retention window.
- Cache read tokens — billed at $0.25/MTok (0.025x input). You pay this on every subsequent request that hits the cached prefix.
- Output tokens — billed at $50/MTok regardless of caching.
The economics reduce to a simple question: does the savings from reading cached tokens (at $0.25 instead of $10 per MTok) exceed the one-time cost of writing them (at $12.50 or $20 per MTok)?
Breakeven analysis: when does a cache write pay for itself?
5-minute TTL (ephemeral)
Writing P tokens to the 5-minute cache costs P × $12.50/MTok. Each subsequent cache read saves P × ($10 - $0.25)/MTok = P × $9.75/MTok.
Breakeven reads = $12.50 / $9.75 ≈ 1.28
After just 2 cache reads, the write has paid for itself. On the old Fable 5 pricing ($1.00 cache reads), breakeven required $12.50 / $9.00 ≈ 1.39 reads — also 2, but with a thinner margin. The Fable 5.1 discount does not change the breakeven count for ephemeral caching, but it dramatically increases the savings per read beyond breakeven.
Savings per additional read after breakeven:
- Fable 5: saves $9.00/MTok per read
- Fable 5.1: saves $9.75/MTok per read (+8.3%)
1-hour TTL (extended)
Writing P tokens to the 1-hour cache costs P × $20/MTok. Each read saves P × $9.75/MTok.
Breakeven reads = $20 / $9.75 ≈ 2.05
You need 3 cache reads within the hour. On Fable 5 (reads at $1.00), breakeven was $20 / $9.00 ≈ 2.22 — also 3 reads. Again, the count is the same, but the per-read savings after breakeven are larger.
The real question: cache-hit ratio
In practice, the breakeven calculation depends on your cache-hit ratio — the fraction of input tokens that come from cache reads rather than fresh input. Anthropic reports that typical workloads achieve 25% effective cost reduction, and agentic workloads reach up to 45%.
Working backward from those figures:
| Scenario | Cache-hit ratio | Effective input cost | Savings vs uncached |
|---|---|---|---|
| No caching | 0% | $10.00/MTok | 0% |
| Light caching (short conversations) | 30-40% | ~$7.00/MTok | ~30% |
| Moderate caching (RAG with stable context) | 50-60% | ~$5.25/MTok | ~48% |
| Heavy caching (long agent runs) | 80-90% | ~$2.25/MTok | ~78% |
At 80% cache-hit ratio, Fable 5.1 input costs drop below Opus 5's uncached rate of $5/MTok. That is the threshold where the cache discount makes a frontier model cheaper on input than a tier below.
Request patterns that benefit most
Agentic loops (highest savings)
An agent sends a growing context on every step: system prompt, tool definitions, task description, and the full transcript of previous steps. The prefix is stable and large. After the first step, nearly every input token is a cache read.
A 50-step agent run with a 50K-token system prompt and growing context can achieve 85-95% cache-hit ratios. On Fable 5.1, that translates to effective input costs between $1.50 and $2.50 per MTok — compared to $10 uncached.
RAG with stable preamble (moderate savings)
If your retrieval-augmented generation pipeline prepends a large, stable system prompt or reference document before the retrieved chunks, the stable prefix caches well. The retrieved chunks and user query change per request, so they remain uncached.
Typical cache-hit ratios: 40-60%, depending on the ratio of stable preamble to dynamic content.
Multi-turn conversations (variable savings)
Each turn appends the previous exchange to the context. The growing prefix caches naturally with automatic caching. Early turns have low cache-hit ratios; later turns approach the agent-loop pattern.
For conversations averaging 8-12 turns, expect 50-70% cache-hit ratios across the session.
Single-shot requests (no savings)
A request that runs once with unique content — a one-off summarization, a single classification — has nothing to cache. If you write to the cache and never read, you paid the write premium (25% or 100% surcharge) for nothing.
System prompt and prefix structuring strategies
Cache hits require an exact prefix match. Structuring your requests to maximize the stable prefix is where engineering effort translates directly into cost savings.
Pattern 1: static system prompt + tools first, user content last
{
"model": "claude-fable-5.1",
"cache_control": {"type": "ephemeral"},
"system": "Your 5,000-token system prompt here...",
"tools": [...],
"messages": [
{"role": "user", "content": "Variable user input"}
]
}
The system prompt and tool definitions form the cached prefix. User content changes per request but does not break the cache for the prefix above it.
Pattern 2: reference documents as early messages
For RAG-style workflows, insert your reference documents as early user/assistant exchanges before the actual query:
{
"model": "claude-fable-5.1",
"cache_control": {"type": "ephemeral"},
"system": "You are a research assistant...",
"messages": [
{"role": "user", "content": "[Reference doc A — 20K tokens]"},
{"role": "assistant", "content": "I've reviewed the reference material."},
{"role": "user", "content": "Based on the above, answer: [variable query]"}
]
}
The system prompt + reference doc form a ~20K-token cached prefix. The variable query is the only uncached input.
Pattern 3: explicit breakpoints for multi-section caching
When you need fine-grained control, place cache_control on specific content blocks:
response = client.messages.create(
model="claude-fable-5.1",
max_tokens=4096,
system=[
{
"type": "text",
"text": "Your stable system prompt...",
"cache_control": {"type": "ephemeral"}
}
],
messages=[
{
"role": "user",
"content": [
{
"type": "text",
"text": large_reference_document,
"cache_control": {"type": "ephemeral"}
},
{
"type": "text",
"text": "Now answer my specific question..."
}
]
}
]
)
This caches both the system prompt and the reference document separately, giving you two cache breakpoints.
Routing decisions: when Fable 5.1 beats a cheaper model
The cache discount creates a non-obvious routing question: when does cached Fable 5.1 cost less per request than an uncached cheaper model?
Fable 5.1 cached vs Opus 5 uncached
| Metric | Fable 5.1 (80% cached) | Opus 5 (uncached) |
|---|---|---|
| Effective input cost | ~$2.20/MTok | $5.00/MTok |
| Output cost | $50/MTok | $25/MTok |
Fable 5.1 wins on input at high cache-hit ratios, but its output cost is 2x Opus 5. For tasks that are input-heavy (large context, short answers), cached Fable 5.1 can be cheaper than uncached Opus 5. For output-heavy tasks (code generation, long-form writing), the output premium dominates.
Fable 5.1 cached vs Sonnet 5 uncached
| Metric | Fable 5.1 (80% cached) | Sonnet 5 (uncached) |
|---|---|---|
| Effective input cost | ~$2.20/MTok | $2.00/MTok |
| Output cost | $50/MTok | $10/MTok |
Even at 80% cache hit, Fable 5.1 input cost only approaches Sonnet 5's uncached input price, while output remains 5x more expensive. Sonnet 5 wins on pure cost unless you specifically need Fable-tier reasoning and the input savings offset the output premium for your workload.
The routing heuristic
Cache-aware routing makes sense when:
- Your workload has a high cache-hit ratio (60%+ of input tokens are cache reads)
- The task is input-heavy (large context window, relatively short outputs)
- You need Fable-tier reasoning (the task fails or degrades on cheaper models)
If all three conditions hold, Fable 5.1 with caching can be the most cost-effective frontier option. If any fails, consider routing to Opus 5 or Sonnet 5 with their respective caching behavior (both use the standard 0.1x cache-read multiplier).
When we route requests through TheRouter, fallback chains can include cache-aware decisions: try Fable 5.1 for high-complexity tasks where the context is largely cached, and fall back to Opus 5 or Sonnet 5 for tasks where caching does not apply or output volume dominates cost.
Comparison: Fable 5.1 caching vs other providers
Prompt caching is not unique to Anthropic. Here is how the Fable 5.1 cache-read discount compares with competing providers:
| Provider | Model | Cache read multiplier | Cache read price (per MTok) | Mechanism |
|---|---|---|---|---|
| Anthropic | Fable 5.1 | 0.025x | $0.25 | Explicit or automatic, 5-min/1-hour TTL |
| Anthropic | Opus 5 | 0.1x | $0.50 | Same mechanism |
| OpenAI | GPT-5.5 Pro | 0.5x | $1.25 | Automatic, no code changes |
| OpenAI | GPT-5.5 Pro (extended 24h) | 0.25x | $0.625 | Automatic, 24h retention |
| DashScope | Qwen3.7-Max | 0.1x (via context cache) | ¥0.1/1K tokens | Explicit context cache API |
Fable 5.1's 0.025x multiplier is the most aggressive cache-read discount available from any major provider on a frontier model. OpenAI's extended 24h retention at 0.25x is competitive in absolute terms, but applies to a different price base.
For a deeper cross-provider caching comparison, see our Prompt Caching Guide: OpenAI, Anthropic, and DashScope.
Common pitfalls
Writing without reading. If you enable caching on requests that rarely repeat, you pay the 25% write surcharge (or 100% for 1-hour TTL) and never recover it. Audit your cache-read/write ratio in the usage response before assuming savings.
Variable content breaking the prefix. A timestamp, request ID, or dynamic value inserted early in the prompt invalidates the cache for everything after it. Move all variable content to the end of the message array.
TTL mismatch. The 5-minute TTL is cheaper to write but expires quickly. If your request cadence is slower than one request per 5 minutes, the cache expires before the second read. Use 1-hour TTL for batch workflows with gaps, and 5-minute for real-time agents.
Ignoring the tokenizer change. Claude 4.7 and later models (including Fable 5.1) use a newer tokenizer that produces approximately 30% more tokens for the same text. Factor this into your cost projections when migrating from older models.
Production checklist
Before deploying Fable 5.1 with caching enabled:
- Measure your cache-hit ratio. Log
cache_read_input_tokensandcache_creation_input_tokensfrom the usage response for at least 24 hours of production traffic. - Calculate your effective input cost. Use the formula:
(uncached_tokens × $10 + cache_write_tokens × $12.50 + cache_read_tokens × $0.25) / total_input_tokens. - Compare against alternatives. If your effective input cost exceeds $5/MTok, consider whether Opus 5 (at $5/MTok uncached) delivers acceptable quality at lower cost.
- Structure prompts for maximum prefix overlap. Static content first, dynamic content last. Use explicit breakpoints if automatic caching is not placing the breakpoint where you want it.
- Set the right TTL. Use 5-minute for high-frequency agent loops. Use 1-hour for batch or lower-frequency workflows.
- Monitor the write/read ratio. A healthy caching deployment should show a write/read ratio well below 1:2. If writes dominate, your cache is thrashing.
- Account for output costs. Fable 5.1 output at $50/MTok is the largest cost component for most workloads. Cache savings on input do not reduce output costs.
FAQ
Does the 75% cache-read discount apply to all Claude models?
No. The 0.025x cache-read multiplier (yielding $0.25/MTok) is exclusive to Fable 5.1 and Mythos 5.1. All other Claude models use a 0.1x multiplier on cache reads. Opus 5 cache reads cost $0.50/MTok; Sonnet 5 cache reads cost $0.20/MTok.
Is prompt caching automatic on Fable 5.1?
Caching is not enabled by default. You must add cache_control to your request — either at the top level for automatic caching or on individual content blocks for explicit breakpoints. Without it, every token is billed at the full input rate.
Does caching affect output quality?
No. Caching stores intermediate computation (attention KV tensors) for the input prefix. The model's generation behavior is identical whether the input was cached or freshly processed.
Can I use prompt caching through TheRouter?
TheRouter routes OpenAI-compatible requests to configured providers. When the upstream provider is Anthropic and the request includes cache_control, the caching behavior passes through to the Anthropic API. TheRouter supports provider/model routing and fallback across configured providers.
How does the newer tokenizer affect caching costs?
Fable 5.1 uses a tokenizer that produces roughly 30% more tokens for the same text compared to Claude 4.6 and earlier. This means the same system prompt costs ~30% more in token terms, but the 75% cache-read discount more than compensates for the tokenizer expansion on cached content.