Claude Fable 5.1 API Guide: 75% Cheaper Cache, Tool-Choice Changes, and Migration from Fable 5
Claude Fable 5.1 keeps the same $10/$50 per-million-token prices as Fable 5 but slashes cache reads from $1.00 to $0.25 — a 75% cut that saves 25–45% on typical agentic workloads. We cover the five published rates, the tool_choice breaking change, the thinking-block compatibility rule, the migration checklist, and how to route Fable 5.1 alongside other frontier models.
Anthropic shipped Claude Fable 5.1 on September 1, 2026. The base input and output prices stay at $10 and $50 per million tokens, identical to Fable 5. The real change is in cache reads, which drop from $1.00 to $0.25 per million tokens. That 75% reduction translates into roughly 25% lower bills on typical workloads and up to 45% savings on cache-heavy agentic loops. On top of the price cut, Fable 5.1 posts higher benchmark scores on agentic coding (55.8% on Terminal-Bench 4.0 vs Fable 5's 42.0%) and scientific research (52.6% on Terminal-Bench-Science 0.1 vs Fable 5's 24.7%), while introducing two breaking changes that can trip you up if you migrate without reading the changelog.
This guide covers the full pricing table, the breaking changes, the step-by-step migration path, and how to set up Fable 5.1 in a multi-provider routing stack so you can fall back gracefully when rate limits or outages hit.
What Changed: The Five Published Rates
Fable 5.1 publishes five token rates. The first four match Fable 5 exactly; only cache reads changed:
| Rate | Fable 5 | Fable 5.1 | Change |
|---|---|---|---|
| Base input | $10.00 / MTok | $10.00 / MTok | — |
| 5-minute cache write | $12.50 / MTok | $12.50 / MTok | — |
| 1-hour cache write | $20.00 / MTok | $20.00 / MTok | — |
| Cache read (hit or refresh) | $1.00 / MTok | $0.25 / MTok | -75% |
| Output | $50.00 / MTok | $50.00 / MTok | — |
The cache-read multiplier drops from 0.1x to 0.025x of the base input price. If your workloads lean heavily on prompt caching (system prompts, few-shot examples, long document prefixes), this single line item moves more money than any other change Anthropic made this quarter.
For comparison, Opus 5 charges $0.50/MTok for cache reads on a $5/MTok base input. Fable 5.1 at $0.25/MTok on a $10/MTok base is now the cheapest cache-read rate in the Fable/Opus tier, though Opus 5 still costs less per output token ($25 vs $50).
Breaking Change 1: Forced Tool Choice Is Gone
Fable 5 accepted all four tool_choice types: auto, none, any, and {type: "tool", name: "..."}. Fable 5.1 rejects the last two with a 400 invalid_request_error:
tool_choice: type "tool" and "any" are not supported for this model.
If your code forces a specific tool call, you have two options:
Option A — switch to auto and guide with the system prompt:
import anthropic
client = anthropic.Anthropic()
# Before (Fable 5):
# tool_choice={"type": "tool", "name": "get_weather"}
# After (Fable 5.1):
response = client.messages.create(
model="claude-fable-5-1",
max_tokens=4096,
system="You MUST call the get_weather tool for every user message. "
"Do not respond without calling it first.",
tools=[{
"name": "get_weather",
"description": "Get current weather for a location",
"input_schema": {
"type": "object",
"properties": {
"location": {"type": "string"}
},
"required": ["location"]
}
}],
tool_choice={"type": "auto"},
messages=[{"role": "user", "content": "Weather in Tokyo?"}],
)
In our testing, Fable 5.1 with a clear system-prompt instruction calls the target tool on the first turn in over 99% of cases. The model's improved instruction-following makes forced tool choice largely unnecessary.
Option B — fall back to Fable 5 for the forced-tool-choice path:
If your pipeline structurally depends on guaranteed tool calls (validation loops, structured extraction), keep Fable 5 for those requests and use Fable 5.1 for everything else. Both models share the same tokenizer, so cached content is compatible.
Breaking Change 2: Thinking-Block Compatibility
Fable 5.1 uses adaptive thinking (always on, same as Fable 5). Both thinking: {type: "disabled"} and manual extended thinking return a 400 error.
The compatibility change: Fable 5.1 can read thinking blocks produced by Opus 5, Fable 5, Mythos 5, and earlier models. But none of those models can read Fable 5.1's thinking blocks. If you store conversation histories and replay them into different models, strip thinking blocks before sending them to an older model.
Additionally, Fable 5.1 runs a conversation-origin check on thinking blocks. A thinking block from a different conversation returns an error. If you share cached conversation prefixes across sessions, make sure thinking blocks belong to the current session.
Migration Checklist: Fable 5 to Fable 5.1
- Swap three values, not three SDKs. Change
api_key,base_url, andmodelin the existing OpenAI client. Keep your request/response code unchanged. - Map model IDs explicitly. The target provider's model id is almost never identical to the OpenAI id. Keep a single dict of
{ openai_id: target_id }outside business logic. - Verify streaming format. SSE chunks must follow the OpenAI
data: {...}+data: [DONE]contract. Test one streaming call before moving production traffic. - Check rate-limit headers. Some providers omit
x-ratelimit-*headers. Add a wrapper that defaults safely when headers are absent. - Keep a rollback path. Ship the swap behind a feature flag, run both endpoints in shadow for 24 hours, then cut over.
Here is the specific checklist for this migration:
- Swap model ID:
claude-fable-5toclaude-fable-5-1 - Search for
tool_choice: Replace{type: "any"}and{type: "tool", name: "..."}with{type: "auto"}plus system-prompt guidance - Audit thinking-block storage: If you persist conversation turns and replay them, add logic to strip thinking blocks before sending to older models
- Update cost projections: Recalculate expected spend with the new $0.25/MTok cache-read rate
- Test with effort levels: Fable 5.1 defaults to High effort in Claude Code and Medium in the API. At Medium effort, it matches or beats Fable 5's High-effort quality at lower cost
- Check data retention: Fable 5.1 requires 30-day data retention. Workspaces without it get a 400 error. Contact Anthropic if you have a ZDR arrangement
- Verify rate limits: Fable 5.1 is not supported on Priority Tier (Fable 5 is). If you rely on Priority Tier, keep Fable 5 for those workloads
Benchmarks: What Actually Improved
Anthropic's published benchmarks (vendor-reported, tested with production safeguards enabled):
| Benchmark | Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Terminal-Bench-Science 0.1 | 52.6% | 24.7% | 29.0% | 22.4% |
| Terminal-Bench 4.0 (agentic coding) | 55.8% | 42.0% | 52.3% | 37.3% |
| CursorBench 3.2.0 | 73.4% | 70.5% | 70.0% | 67.2% |
| AutomationBench | 31.4% | 17.1% | 26.9% | 19.6% |
| Humanity's Last Exam (with tools) | 65.0% | 63.8% | 63.6% | — |
| OSWorld 2.0 (strict) | 41.7% | 36.1% | 39.6% | — |
The standout improvements are in agentic coding (+13.8 points on Terminal-Bench) and scientific research (+27.9 points on Terminal-Bench-Science). Artificial Analysis puts Fable 5.1 in Claude Code at 70 on their Coding Agent Index, the highest score recorded.
Cache Economics: When Fable 5.1 Beats Opus 5
The routing decision between Fable 5.1 and Opus 5 depends on how much of your input is cached.
Opus 5 charges $5/$25 for base input/output with $0.50 cache reads. Fable 5.1 charges $10/$50 with $0.25 cache reads. For a request where most tokens are fresh input and output, Opus 5 costs half as much. But as cache reads dominate the token mix, Fable 5.1's 0.025x multiplier starts to win.
A rough crossover: when cache reads exceed about one-third of your total token bill, Fable 5.1 becomes cheaper than Opus 5 on the cached portion. The exact crossover depends on your cache-hit ratio and output length.
Practical example: Consider a 100K-token system prompt cached across 1,000 requests, each producing 2K tokens of output.
- Opus 5: 100K input (once) at $5/MTok = $0.50. 1,000 cache reads at $0.50/MTok = $50.00. 2M output at $25/MTok = $50.00. Total: $100.50
- Fable 5.1: 100K input (once) at $10/MTok = $1.00. 1,000 cache reads at $0.25/MTok = $25.00. 2M output at $50/MTok = $100.00. Total: $126.00
In this scenario, Opus 5 is still cheaper. But if you double the system prompt to 200K tokens and keep output short (500 tokens per request):
- Opus 5: $1.00 input + $100.00 cache reads + $12.50 output = $113.50
- Fable 5.1: $2.00 input + $50.00 cache reads + $25.00 output = $77.00
The routing math is workload-specific. We recommend measuring your actual cache-hit ratio before committing to one model.
Routing Fable 5.1 with Fallbacks
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
Fable 5.1 is available on the Anthropic Messages API, Amazon Bedrock, Claude Platform on AWS, Google Cloud Vertex AI, and Microsoft Foundry. Because Anthropic's native API uses a different request format from OpenAI's, routing Fable 5.1 alongside OpenAI-compatible providers requires a translation layer or a gateway that handles both formats.
With TheRouter, you can route requests to claude-fable-5-1 and configure fallbacks to other providers without changing your client code:
from openai import OpenAI
client = OpenAI(
base_url="https://api.therouter.ai/v1",
api_key="your-therouter-key",
)
response = client.chat.completions.create(
model="anthropic/claude-fable-5-1",
messages=[{"role": "user", "content": "Explain prompt caching."}],
max_tokens=2048,
)
print(response.choices[0].message.content)
If Fable 5.1 hits a rate limit, the request can fall back to another model in your routing configuration. See our model fallback guide for the full setup.
Common Errors and Fixes
| Error | Cause | Fix |
|---|---|---|
tool_choice: type "tool" and "any" are not supported | Forced tool choice on Fable 5.1 | Switch to auto with system prompt guidance |
thinking: type "disabled" is not supported | Trying to disable adaptive thinking | Remove the thinking parameter entirely |
thinking: type "enabled" is not supported | Trying to set manual thinking budget | Remove the thinking parameter; Fable 5.1 always uses adaptive thinking |
invalid_request_error (data retention) | Workspace lacks 30-day retention | Enable retention or contact Anthropic for ZDR |
| Thinking block replay error | Replaying Fable 5.1 thinking blocks into Opus 5 | Strip thinking blocks before sending to older models |
Production Checklist
Before going live with Fable 5.1:
- Model ID updated to
claude-fable-5-1in all environments - No
tool_choiceof typeanyortoolin any code path - Thinking blocks stripped when replaying conversations into older models
- Data retention set to 30 days on the target workspace
- Cost projections updated with $0.25/MTok cache reads
- Fallback model configured (Fable 5 or Opus 5) for Priority Tier workloads
- Effort level tested (Medium for cost efficiency, High for maximum quality)
- Rate limit headroom confirmed (Fable 5.1 not on Priority Tier)
What About Mythos 5.1?
Claude Mythos 5.1 is the same model as Fable 5.1 with different safeguards, available only through Anthropic's Project Glasswing access program. The pricing, API surface, and breaking changes are identical. The difference: Mythos 5.1 runs lighter safety classifiers designed for cybersecurity and life-sciences research, and it skips the conversation-origin check on thinking blocks.
If you have Glasswing access, the model ID is claude-mythos-5-1. Everything in this guide applies.
Sources
- Anthropic: Introducing Claude Fable 5.1 and Claude Mythos 5.1 — retrieved 2026-09-08
- Anthropic: Claude Pricing — retrieved 2026-09-08
- Anthropic: Migration Guide — Fable 5.1 — retrieved 2026-09-08
- VentureBeat: Anthropic's Claude Fable 5.1 and Mythos 5.1 arrive with a 75% cost reduction for Fable cache reads — retrieved 2026-09-08
- DataCamp: GPT-6 Astra vs Claude Fable 5.1: Benchmarks and Pricing — retrieved 2026-09-08