Claude Sonnet 5's Three Breaking API Changes: What Every Operator Must Audit Before Migrating

Claude Sonnet 5 ships three silent production killers: adaptive thinking on by default, temperature/top_p/top_k now returns 400, and a new tokenizer that inflates token counts ~30%. Here's the operator checklist.

TheRouter Newsroomvia Anthropic
Technical diagram showing Claude Sonnet 5 API parameter changes with 400 error paths and tokenizer delta annotations

The Sonnet 5 launch coverage focused on pricing and benchmark numbers. What it underplayed: three API-level behavioral changes that will silently break production integrations for teams migrating from Sonnet 4.6. They are not optional edge cases—they fire on the most common API call patterns.

If your routing policy is already pointing traffic at claude-sonnet-5, run the audit below now.

What changed

Anthropic's official What's new in Claude Sonnet 5 page documents three breaking changes and one silent cost multiplier:

1. Adaptive thinking is on by default

On Sonnet 4.6, requests without a thinking field ran without any extended thinking. On Sonnet 5, the same requests now run with adaptive thinking enabled. The model decides when and how much to think on each request.

To disable thinking explicitly:

thinking = {"type": "disabled"}

The impact is twofold. First, any workload that relied on no-thinking latency profiles from Sonnet 4.6 now has unpredictable latency unless thinking is explicitly disabled. Second, because max_tokens is a hard cap on total output—thinking tokens included—a max_tokens value tuned for Sonnet 4.6 response text may now truncate output once thinking tokens are factored in. Revisit all max_tokens limits on workloads that did not previously use thinking.

2. Sampling parameters return 400

Setting temperature, top_p, or top_k to any non-default value now returns a 400 error. This is the same constraint Anthropic introduced on Opus 4.7 and Fable 5; it now applies to the Sonnet tier for the first time.

# Will return 400 on Sonnet 5
response = client.messages.create(
    model="claude-sonnet-5",
    temperature=0.7,  # NOT ALLOWED
    ...
)

# Remove the parameter or set the default explicitly
response = client.messages.create(
    model="claude-sonnet-5",
    # temperature omitted — uses default
    ...
)

Any system that passes non-default temperature, top_p, or top_k values to Sonnet 4.6 will receive a hard 400 error after migration. This covers a very large share of production integrations: LangChain, LlamaIndex, and most SDK wrappers expose temperature as a top-level parameter with non-zero defaults.

3. Manual extended thinking removed

The thinking: {type: "enabled", budget_tokens: N} pattern was deprecated on Sonnet 4.6. On Sonnet 5 it is removed entirely and returns a 400 error.

# Not supported on Sonnet 5 (400 error)
thinking = {"type": "enabled", "budget_tokens": 32000}

# Use adaptive instead
thinking = {"type": "adaptive"}
# or explicitly set effort level
effort = {"level": "high"}

If you have production code that calls Sonnet 4.6 with a manual thinking budget—for example, a reasoning-intensive code-review workflow—that code will break on Sonnet 5 without modification.

4. New tokenizer — ~30% more tokens for the same text

Sonnet 5 uses a new tokenizer. The same input text produces approximately 30% more tokens than on Sonnet 4.6. This is not an API contract change—the request/response shape is identical—but it affects everything you measure or budget in tokens:

  • usage fields: token counts in API responses will be higher for equivalent prompts.
  • Context window capacity: the 1M-token window holds less text per token, so prompts close to the old context limit may now overflow.
  • max_tokens output budgets: output limits sized for Sonnet 4.6 may truncate shorter responses on Sonnet 5.
  • Per-request cost: input/output pricing is unchanged per token, but each token now covers less text. Cost for an equivalent prompt can be up to 30% higher.

Do not reuse token counts measured against Sonnet 4.6. Run token-counting calls against Sonnet 5 directly for any budget that matters.

Why it matters for AI engineering teams

Teams using an AI gateway or multi-provider routing layer need to think about these changes at two levels.

For direct API callers the risks are clear: hard 400 errors on temperature and thinking parameters if not removed before migration; silent latency degradation if adaptive thinking fires on low-latency paths; silent cost increases if token budgets are not recalibrated.

For gateway operators and routing policies, the concern is subtler. If you route claude-sonnet-latest to Sonnet 5 as part of an automatic alias update, any downstream caller still passing temperature will immediately start receiving 400s. If your gateway proxies parameter fields through without scrubbing model-incompatible parameters, you become the silent failure point. This is an argument for routing policies that include parameter transformation rules, not just model name substitution.

The same problem applies to fallback chains. If your primary model is GPT-5.5 or Gemini 3.5 Flash with temperature=0.8 and your fallback is Sonnet 5, every fallback invocation will return a 400 instead of a successful fallback response.

The router/operator angle

Parameter compatibility matrix for routing tiers

Maintaining a routing policy across providers requires knowing which parameters each model accepts. Sonnet 5 raises the stakes:

ParameterSonnet 4.6Sonnet 5
temperature (non-default)✓ Accepted✗ 400 error
top_p (non-default)✓ Accepted✗ 400 error
top_k (non-default)✓ Accepted✗ 400 error
thinking.type: "enabled"Deprecated✗ 400 error
thinking.type: "adaptive"✓ Accepted✓ Accepted (default)
thinking.type: "disabled"✓ Accepted✓ Accepted

Gateway-level parameter normalization—stripping or defaulting model-incompatible fields before forwarding—is now a correctness requirement, not just a nicety, if you route traffic to Sonnet 5 from callers that set temperature.

Token budget recalibration checklist

Before switching a production routing tier to Sonnet 5:

  1. Re-run token counting on your 10th, 50th, and 90th percentile prompt lengths against claude-sonnet-5. The baseline has shifted ~30%.
  2. Raise max_tokens budgets on output-heavy paths by at least 30% to avoid truncation.
  3. Re-check context-window utilization on workloads that run near the limit. Prompts that fit comfortably on Sonnet 4.6 may overflow on Sonnet 5 if they exceed ~770K Sonnet-4.6 tokens worth of text.
  4. Re-run cost projections. Recompute expected monthly cost for your top request types against Sonnet 5 token counts before the August 31 pricing window closes.

Thinking token accounting on routing layers

Adaptive thinking adds hidden token consumption. If your routing layer bills downstream users based on token counts from the usage field, thinking tokens now appear in output token totals. This changes the billing semantics of your routing tier. Sonnet 5's usage.output_tokens includes thinking tokens when adaptive thinking fires; Sonnet 4.6's did not (for the same calls without thinking enabled).

If you surface per-request cost to your users, you likely need to handle thinking token accounting separately or document the change.

What TheRouter users should watch or try

Teams routing through TheRouter or any AI gateway to Anthropic should:

  1. Check whether your provider's Sonnet alias points to claude-sonnet-5 now or will update automatically. Confirm with your gateway's model list before the switch happens.
  2. If routing to claude-sonnet-5 directly, run a parameter compatibility check against your request schema—specifically for temperature, top_p, top_k, and thinking fields.
  3. Review the Claude API models overview for current model IDs and the Anthropic migration guide for a structured migration checklist.

The /claude-api migrate skill in Claude Code can also automate parameter fixes across a codebase if you need to update multiple integration points at once.

Models covered in this article

Help & contact