Claude Sonnet 5's Three Breaking API Changes: What Every Operator Must Audit Before Migrating
Claude Sonnet 5 ships three silent production killers: adaptive thinking on by default, temperature/top_p/top_k now returns 400, and a new tokenizer that inflates token counts ~30%. Here's the operator checklist.

The Sonnet 5 launch coverage focused on pricing and benchmark numbers. What it underplayed: three API-level behavioral changes that will silently break production integrations for teams migrating from Sonnet 4.6. They are not optional edge cases—they fire on the most common API call patterns.
If your routing policy is already pointing traffic at claude-sonnet-5, run the audit below now.
What changed
Anthropic's official What's new in Claude Sonnet 5 page documents three breaking changes and one silent cost multiplier:
1. Adaptive thinking is on by default
On Sonnet 4.6, requests without a thinking field ran without any extended thinking. On Sonnet 5, the same requests now run with adaptive thinking enabled. The model decides when and how much to think on each request.
To disable thinking explicitly:
thinking = {"type": "disabled"}
The impact is twofold. First, any workload that relied on no-thinking latency profiles from Sonnet 4.6 now has unpredictable latency unless thinking is explicitly disabled. Second, because max_tokens is a hard cap on total output—thinking tokens included—a max_tokens value tuned for Sonnet 4.6 response text may now truncate output once thinking tokens are factored in. Revisit all max_tokens limits on workloads that did not previously use thinking.
2. Sampling parameters return 400
Setting temperature, top_p, or top_k to any non-default value now returns a 400 error. This is the same constraint Anthropic introduced on Opus 4.7 and Fable 5; it now applies to the Sonnet tier for the first time.
# Will return 400 on Sonnet 5
response = client.messages.create(
model="claude-sonnet-5",
temperature=0.7, # NOT ALLOWED
...
)
# Remove the parameter or set the default explicitly
response = client.messages.create(
model="claude-sonnet-5",
# temperature omitted — uses default
...
)
Any system that passes non-default temperature, top_p, or top_k values to Sonnet 4.6 will receive a hard 400 error after migration. This covers a very large share of production integrations: LangChain, LlamaIndex, and most SDK wrappers expose temperature as a top-level parameter with non-zero defaults.
3. Manual extended thinking removed
The thinking: {type: "enabled", budget_tokens: N} pattern was deprecated on Sonnet 4.6. On Sonnet 5 it is removed entirely and returns a 400 error.
# Not supported on Sonnet 5 (400 error)
thinking = {"type": "enabled", "budget_tokens": 32000}
# Use adaptive instead
thinking = {"type": "adaptive"}
# or explicitly set effort level
effort = {"level": "high"}
If you have production code that calls Sonnet 4.6 with a manual thinking budget—for example, a reasoning-intensive code-review workflow—that code will break on Sonnet 5 without modification.
4. New tokenizer — ~30% more tokens for the same text
Sonnet 5 uses a new tokenizer. The same input text produces approximately 30% more tokens than on Sonnet 4.6. This is not an API contract change—the request/response shape is identical—but it affects everything you measure or budget in tokens:
usagefields: token counts in API responses will be higher for equivalent prompts.- Context window capacity: the 1M-token window holds less text per token, so prompts close to the old context limit may now overflow.
max_tokensoutput budgets: output limits sized for Sonnet 4.6 may truncate shorter responses on Sonnet 5.- Per-request cost: input/output pricing is unchanged per token, but each token now covers less text. Cost for an equivalent prompt can be up to 30% higher.
Do not reuse token counts measured against Sonnet 4.6. Run token-counting calls against Sonnet 5 directly for any budget that matters.
Why it matters for AI engineering teams
Teams using an AI gateway or multi-provider routing layer need to think about these changes at two levels.
For direct API callers the risks are clear: hard 400 errors on temperature and thinking parameters if not removed before migration; silent latency degradation if adaptive thinking fires on low-latency paths; silent cost increases if token budgets are not recalibrated.
For gateway operators and routing policies, the concern is subtler. If you route claude-sonnet-latest to Sonnet 5 as part of an automatic alias update, any downstream caller still passing temperature will immediately start receiving 400s. If your gateway proxies parameter fields through without scrubbing model-incompatible parameters, you become the silent failure point. This is an argument for routing policies that include parameter transformation rules, not just model name substitution.
The same problem applies to fallback chains. If your primary model is GPT-5.5 or Gemini 3.5 Flash with temperature=0.8 and your fallback is Sonnet 5, every fallback invocation will return a 400 instead of a successful fallback response.
The router/operator angle
Parameter compatibility matrix for routing tiers
Maintaining a routing policy across providers requires knowing which parameters each model accepts. Sonnet 5 raises the stakes:
| Parameter | Sonnet 4.6 | Sonnet 5 |
|---|---|---|
temperature (non-default) | ✓ Accepted | ✗ 400 error |
top_p (non-default) | ✓ Accepted | ✗ 400 error |
top_k (non-default) | ✓ Accepted | ✗ 400 error |
thinking.type: "enabled" | Deprecated | ✗ 400 error |
thinking.type: "adaptive" | ✓ Accepted | ✓ Accepted (default) |
thinking.type: "disabled" | ✓ Accepted | ✓ Accepted |
Gateway-level parameter normalization—stripping or defaulting model-incompatible fields before forwarding—is now a correctness requirement, not just a nicety, if you route traffic to Sonnet 5 from callers that set temperature.
Token budget recalibration checklist
Before switching a production routing tier to Sonnet 5:
- Re-run token counting on your 10th, 50th, and 90th percentile prompt lengths against
claude-sonnet-5. The baseline has shifted ~30%. - Raise
max_tokensbudgets on output-heavy paths by at least 30% to avoid truncation. - Re-check context-window utilization on workloads that run near the limit. Prompts that fit comfortably on Sonnet 4.6 may overflow on Sonnet 5 if they exceed ~770K Sonnet-4.6 tokens worth of text.
- Re-run cost projections. Recompute expected monthly cost for your top request types against Sonnet 5 token counts before the August 31 pricing window closes.
Thinking token accounting on routing layers
Adaptive thinking adds hidden token consumption. If your routing layer bills downstream users based on token counts from the usage field, thinking tokens now appear in output token totals. This changes the billing semantics of your routing tier. Sonnet 5's usage.output_tokens includes thinking tokens when adaptive thinking fires; Sonnet 4.6's did not (for the same calls without thinking enabled).
If you surface per-request cost to your users, you likely need to handle thinking token accounting separately or document the change.
What TheRouter users should watch or try
Teams routing through TheRouter or any AI gateway to Anthropic should:
- Check whether your provider's Sonnet alias points to
claude-sonnet-5now or will update automatically. Confirm with your gateway's model list before the switch happens. - If routing to
claude-sonnet-5directly, run a parameter compatibility check against your request schema—specifically fortemperature,top_p,top_k, andthinkingfields. - Review the Claude API models overview for current model IDs and the Anthropic migration guide for a structured migration checklist.
The /claude-api migrate skill in Claude Code can also automate parameter fixes across a codebase if you need to update multiple integration points at once.
Models covered in this article

Claude Opus 5.5: Four Breaking API Changes and What They Mean for Your Routing Setup
Four breaking changes in Claude Opus 5.5: thinking can't be disabled, forced tool_choice returns 400, thinking blocks don't cross non-Fable/Mythos models, and computer_20251124 is gone. Each has a specific fix — three carry fallback routing implications the announcement skips.

Claude Platform on AWS: The Third Deployment Path Every Operator Routing Policy Must Now Account For
Claude Platform on AWS gives teams Anthropic-managed inference through AWS billing — full beta headers, Agent Skills, and a separate capacity pool that unlocks a new multi-platform failover strategy.

Claude Opus 4.1 Deprecation: Anthropic August 5 Migration Guide for Router Teams
Anthropic's Claude Opus 4.1 deprecation retires claude-opus-4-1-20250805 on August 5, 2026. Use this router-focused migration guide to replace aliases, audit fallback tiers, and avoid API failures.