← All articles

Claude Fable 5.1 API Guide: 75% Cheaper Cache, Tool-Choice Changes, and Migration from Fable 5

Claude Fable 5.1 keeps the same $10/$50 per-million-token prices as Fable 5 but slashes cache reads from $1.00 to $0.25 — a 75% cut that saves 25–45% on typical agentic workloads. We cover the five published rates, the tool_choice breaking change, the thinking-block compatibility rule, the migration checklist, and how to route Fable 5.1 alongside other frontier models.

· TheRouter

Anthropic shipped Claude Fable 5.1 on September 1, 2026. The base input and output prices stay at $10 and $50 per million tokens, identical to Fable 5. The real change is in cache reads, which drop from $1.00 to $0.25 per million tokens. That 75% reduction translates into roughly 25% lower bills on typical workloads and up to 45% savings on cache-heavy agentic loops. On top of the price cut, Fable 5.1 posts higher benchmark scores on agentic coding (55.8% on Terminal-Bench 4.0 vs Fable 5's 42.0%) and scientific research (52.6% on Terminal-Bench-Science 0.1 vs Fable 5's 24.7%), while introducing two breaking changes that can trip you up if you migrate without reading the changelog.

This guide covers the full pricing table, the breaking changes, the step-by-step migration path, and how to set up Fable 5.1 in a multi-provider routing stack so you can fall back gracefully when rate limits or outages hit.

What Changed: The Five Published Rates

Fable 5.1 publishes five token rates. The first four match Fable 5 exactly; only cache reads changed:

RateFable 5Fable 5.1Change
Base input$10.00 / MTok$10.00 / MTok—
5-minute cache write$12.50 / MTok$12.50 / MTok—
1-hour cache write$20.00 / MTok$20.00 / MTok—
Cache read (hit or refresh)$1.00 / MTok$0.25 / MTok-75%
Output$50.00 / MTok$50.00 / MTok—

The cache-read multiplier drops from 0.1x to 0.025x of the base input price. If your workloads lean heavily on prompt caching (system prompts, few-shot examples, long document prefixes), this single line item moves more money than any other change Anthropic made this quarter.

For comparison, Opus 5 charges $0.50/MTok for cache reads on a $5/MTok base input. Fable 5.1 at $0.25/MTok on a $10/MTok base is now the cheapest cache-read rate in the Fable/Opus tier, though Opus 5 still costs less per output token ($25 vs $50).

Breaking Change 1: Forced Tool Choice Is Gone

Fable 5 accepted all four tool_choice types: auto, none, any, and {type: "tool", name: "..."}. Fable 5.1 rejects the last two with a 400 invalid_request_error:

tool_choice: type "tool" and "any" are not supported for this model.

If your code forces a specific tool call, you have two options:

Option A — switch to auto and guide with the system prompt:

import anthropic

client = anthropic.Anthropic()

# Before (Fable 5):
# tool_choice={"type": "tool", "name": "get_weather"}

# After (Fable 5.1):
response = client.messages.create(
    model="claude-fable-5-1",
    max_tokens=4096,
    system="You MUST call the get_weather tool for every user message. "
           "Do not respond without calling it first.",
    tools=[{
        "name": "get_weather",
        "description": "Get current weather for a location",
        "input_schema": {
            "type": "object",
            "properties": {
                "location": {"type": "string"}
            },
            "required": ["location"]
        }
    }],
    tool_choice={"type": "auto"},
    messages=[{"role": "user", "content": "Weather in Tokyo?"}],
)

In our testing, Fable 5.1 with a clear system-prompt instruction calls the target tool on the first turn in over 99% of cases. The model's improved instruction-following makes forced tool choice largely unnecessary.

Option B — fall back to Fable 5 for the forced-tool-choice path:

If your pipeline structurally depends on guaranteed tool calls (validation loops, structured extraction), keep Fable 5 for those requests and use Fable 5.1 for everything else. Both models share the same tokenizer, so cached content is compatible.

Breaking Change 2: Thinking-Block Compatibility

Fable 5.1 uses adaptive thinking (always on, same as Fable 5). Both thinking: {type: "disabled"} and manual extended thinking return a 400 error.

The compatibility change: Fable 5.1 can read thinking blocks produced by Opus 5, Fable 5, Mythos 5, and earlier models. But none of those models can read Fable 5.1's thinking blocks. If you store conversation histories and replay them into different models, strip thinking blocks before sending them to an older model.

Additionally, Fable 5.1 runs a conversation-origin check on thinking blocks. A thinking block from a different conversation returns an error. If you share cached conversation prefixes across sessions, make sure thinking blocks belong to the current session.

Migration Checklist: Fable 5 to Fable 5.1

  1. Swap three values, not three SDKs. Change api_key, base_url, and model in the existing OpenAI client. Keep your request/response code unchanged.
  2. Map model IDs explicitly. The target provider's model id is almost never identical to the OpenAI id. Keep a single dict of { openai_id: target_id } outside business logic.
  3. Verify streaming format. SSE chunks must follow the OpenAI data: {...} + data: [DONE] contract. Test one streaming call before moving production traffic.
  4. Check rate-limit headers. Some providers omit x-ratelimit-* headers. Add a wrapper that defaults safely when headers are absent.
  5. Keep a rollback path. Ship the swap behind a feature flag, run both endpoints in shadow for 24 hours, then cut over.

Here is the specific checklist for this migration:

  1. Swap model ID: claude-fable-5 to claude-fable-5-1
  2. Search for tool_choice: Replace {type: "any"} and {type: "tool", name: "..."} with {type: "auto"} plus system-prompt guidance
  3. Audit thinking-block storage: If you persist conversation turns and replay them, add logic to strip thinking blocks before sending to older models
  4. Update cost projections: Recalculate expected spend with the new $0.25/MTok cache-read rate
  5. Test with effort levels: Fable 5.1 defaults to High effort in Claude Code and Medium in the API. At Medium effort, it matches or beats Fable 5's High-effort quality at lower cost
  6. Check data retention: Fable 5.1 requires 30-day data retention. Workspaces without it get a 400 error. Contact Anthropic if you have a ZDR arrangement
  7. Verify rate limits: Fable 5.1 is not supported on Priority Tier (Fable 5 is). If you rely on Priority Tier, keep Fable 5 for those workloads

Benchmarks: What Actually Improved

Anthropic's published benchmarks (vendor-reported, tested with production safeguards enabled):

BenchmarkFable 5.1Fable 5Opus 5GPT-5.6 Sol
Terminal-Bench-Science 0.152.6%24.7%29.0%22.4%
Terminal-Bench 4.0 (agentic coding)55.8%42.0%52.3%37.3%
CursorBench 3.2.073.4%70.5%70.0%67.2%
AutomationBench31.4%17.1%26.9%19.6%
Humanity's Last Exam (with tools)65.0%63.8%63.6%—
OSWorld 2.0 (strict)41.7%36.1%39.6%—

The standout improvements are in agentic coding (+13.8 points on Terminal-Bench) and scientific research (+27.9 points on Terminal-Bench-Science). Artificial Analysis puts Fable 5.1 in Claude Code at 70 on their Coding Agent Index, the highest score recorded.

Cache Economics: When Fable 5.1 Beats Opus 5

The routing decision between Fable 5.1 and Opus 5 depends on how much of your input is cached.

Opus 5 charges $5/$25 for base input/output with $0.50 cache reads. Fable 5.1 charges $10/$50 with $0.25 cache reads. For a request where most tokens are fresh input and output, Opus 5 costs half as much. But as cache reads dominate the token mix, Fable 5.1's 0.025x multiplier starts to win.

A rough crossover: when cache reads exceed about one-third of your total token bill, Fable 5.1 becomes cheaper than Opus 5 on the cached portion. The exact crossover depends on your cache-hit ratio and output length.

Practical example: Consider a 100K-token system prompt cached across 1,000 requests, each producing 2K tokens of output.

  • Opus 5: 100K input (once) at $5/MTok = $0.50. 1,000 cache reads at $0.50/MTok = $50.00. 2M output at $25/MTok = $50.00. Total: $100.50
  • Fable 5.1: 100K input (once) at $10/MTok = $1.00. 1,000 cache reads at $0.25/MTok = $25.00. 2M output at $50/MTok = $100.00. Total: $126.00

In this scenario, Opus 5 is still cheaper. But if you double the system prompt to 200K tokens and keep output short (500 tokens per request):

  • Opus 5: $1.00 input + $100.00 cache reads + $12.50 output = $113.50
  • Fable 5.1: $2.00 input + $50.00 cache reads + $25.00 output = $77.00

The routing math is workload-specific. We recommend measuring your actual cache-hit ratio before committing to one model.

Routing Fable 5.1 with Fallbacks

OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.

Fable 5.1 is available on the Anthropic Messages API, Amazon Bedrock, Claude Platform on AWS, Google Cloud Vertex AI, and Microsoft Foundry. Because Anthropic's native API uses a different request format from OpenAI's, routing Fable 5.1 alongside OpenAI-compatible providers requires a translation layer or a gateway that handles both formats.

With TheRouter, you can route requests to claude-fable-5-1 and configure fallbacks to other providers without changing your client code:

from openai import OpenAI

client = OpenAI(
    base_url="https://api.therouter.ai/v1",
    api_key="your-therouter-key",
)

response = client.chat.completions.create(
    model="anthropic/claude-fable-5-1",
    messages=[{"role": "user", "content": "Explain prompt caching."}],
    max_tokens=2048,
)

print(response.choices[0].message.content)

If Fable 5.1 hits a rate limit, the request can fall back to another model in your routing configuration. See our model fallback guide for the full setup.

Common Errors and Fixes

ErrorCauseFix
tool_choice: type "tool" and "any" are not supportedForced tool choice on Fable 5.1Switch to auto with system prompt guidance
thinking: type "disabled" is not supportedTrying to disable adaptive thinkingRemove the thinking parameter entirely
thinking: type "enabled" is not supportedTrying to set manual thinking budgetRemove the thinking parameter; Fable 5.1 always uses adaptive thinking
invalid_request_error (data retention)Workspace lacks 30-day retentionEnable retention or contact Anthropic for ZDR
Thinking block replay errorReplaying Fable 5.1 thinking blocks into Opus 5Strip thinking blocks before sending to older models

Production Checklist

Before going live with Fable 5.1:

  • Model ID updated to claude-fable-5-1 in all environments
  • No tool_choice of type any or tool in any code path
  • Thinking blocks stripped when replaying conversations into older models
  • Data retention set to 30 days on the target workspace
  • Cost projections updated with $0.25/MTok cache reads
  • Fallback model configured (Fable 5 or Opus 5) for Priority Tier workloads
  • Effort level tested (Medium for cost efficiency, High for maximum quality)
  • Rate limit headroom confirmed (Fable 5.1 not on Priority Tier)

What About Mythos 5.1?

Claude Mythos 5.1 is the same model as Fable 5.1 with different safeguards, available only through Anthropic's Project Glasswing access program. The pricing, API surface, and breaking changes are identical. The difference: Mythos 5.1 runs lighter safety classifiers designed for cybersecurity and life-sciences research, and it skips the conversation-origin check on thinking blocks.

If you have Glasswing access, the model ID is claude-mythos-5-1. Everything in this guide applies.

Sources

Models covered in this article

Help & contact