Claude Sonnet 5 Skips Priority Tier: What the Service-Tier Gap Means for Latency-Sensitive Routing Teams
Claude Sonnet 5 is excluded from Priority Tier — and Priority Tier is now closed to new purchases. Teams with latency-critical workloads face a routing fork: keep Priority Tier on older models or migrate to Sonnet 5 at standard-tier SLAs.
Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

Anthropic's Claude Sonnet 5 launch on June 30 generated a wave of coverage on benchmark improvements and the new tokenizer. What most routing teams missed: Sonnet 5 is explicitly excluded from Priority Tier — and Priority Tier is no longer available for new purchases. For teams that rely on Priority Tier capacity commitments to hit latency SLAs, this creates a routing decision that cannot be deferred past August 31.
What happened
Anthropic's service tiers documentation now carries a warning that "Priority Tier capacity commitments are no longer available for purchase." Organizations with existing commitments can continue using Priority Tier through their contract end date. New entrants to the Claude API cannot acquire Priority Tier access at all.
More immediately, the supported-models table under Existing Priority Tier commitments lists the exclusions explicitly:
Priority Tier is supported on all available Claude models (including Claude Fable 5 and Claude Opus 4.8) except Claude Sonnet 5, Claude Mythos Preview, and Batch tier models.
This means a team currently routing latency-critical traffic through claude-sonnet-4-6 on a Priority Tier commitment cannot simply swap in claude-sonnet-5 and retain their tier protection. The Sonnet 5 request will fall through to standard tier regardless of the committed capacity remaining.
Why it matters for AI engineering teams
Priority Tier was the only mechanism Anthropic offered for reducing "server overloaded" errors during peak demand. The tier works by reserving input and output token capacity per minute. When your committed capacity is available, requests get prioritized over standard-tier traffic. When it runs out, requests fall back to standard tier automatically.
Without Priority Tier, your fallback options for Sonnet 5 traffic are:
- Accept standard-tier SLAs — best-effort availability, no overload protection during traffic spikes.
- Route latency-sensitive requests to Opus 4.8 or Fable 5 on your existing Priority Tier commitment, while routing lower-urgency Sonnet 5 traffic at standard.
- Use multi-provider fallback — configure a secondary provider route so that Sonnet 5 overloaded errors fail over to an alternative endpoint.
Option 3 is the cleanest path for teams that cannot acquire new Priority Tier capacity, and it is exactly the kind of routing policy a gateway layer is built to enforce.
The router/operator angle
The Priority Tier gap compounds with two other Sonnet 5 constraints that operators need to model together:
1. The tokenizer multiplier. Sonnet 5 uses a new tokenizer that produces approximately 30% more tokens for the same input text. Per-token pricing is unchanged, but per-request cost is not — the same prompt costs more tokens to process. Our earlier breakdown covers the full impact; the short version is that cost budgets calibrated against Sonnet 4.6 may underestimate Sonnet 5 spend by 20–35% depending on workload.
2. The August 31 pricing cliff. Anthropic set introductory pricing for Sonnet 5 at $2 / $10 per million tokens (input / output) through August 31, 2026. Starting September 1, standard pricing kicks in at $3 / $15. If your routing layer selects models based on cost thresholds, Sonnet 5 will appear cheaper than it will actually be for any workload that runs past the summer.
The three constraints — no Priority Tier, ~30% token inflation, and a September pricing step-up — mean the true cost-and-latency profile of a Sonnet 5 migration will look different in Q4 than it does today. Teams doing any routing policy review before that date need to stress-test the September pricing in their model.
A routing framework for the Priority Tier gap
For teams evaluating how to handle the Sonnet 5 Priority Tier exclusion, here is a practical decision tree:
If you have an existing Priority Tier commitment:
- Continue routing latency-critical workloads to models covered by the commitment (Fable 5, Opus 4.8, Sonnet 4.6).
- Route Sonnet 5 at standard tier for workloads where occasional overload retries are acceptable.
- Do not assume your committed capacity covers Sonnet 5 — validate the exclusion in your API responses by checking the
usage.service_tierfield.
If you do not have a Priority Tier commitment:
- Priority Tier is closed to new purchases. Your only path to overload resilience is fallback routing.
- Design your Sonnet 5 routing to fail over to a secondary provider or model on
overloadederrors. - Monitor
anthropic-priority-input-tokens-remainingresponse headers — their absence on a Sonnet 5 request confirms you are on standard tier.
Before August 31:
- Lock in your Sonnet 5 routing policy under introductory pricing. Any routing rules that weigh Sonnet 5 as cost-equivalent to Sonnet 4.6 will be incorrect after September 1.
- Account for the tokenizer multiplier when setting
max_tokenslimits and per-request cost ceilings.
What TheRouter users should watch
If you are routing Sonnet 5 traffic through a model gateway, audit the service_tier field in API responses to confirm your Sonnet 5 requests are handled correctly. A service_tier: "standard" response is expected and correct for Sonnet 5 — any routing policy that expects "priority" on a Sonnet 5 request should be updated now.
For fallback configuration, adding a secondary provider path for Sonnet 5 overloaded responses is the closest available substitute for Priority Tier capacity protection. Visit the TheRouter docs for guidance on configuring model fallback chains.
Models covered in this article

Fable 5.1's Cache Read Cut and Effort Parameter Are the Two Changes Every Gateway Operator Must Price In
Cache reads on claude-fable-5-1 drop to $0.25/MTok — 4x cheaper than other Claude models — and output_config.effort lets operators trade thinking depth for token spend per request. What both changes mean for routing, billing, and upgrade decisions.

Anthropic Model Hardware Standard: The New Safety Boundary for Physical AI Agents
Anthropic Model Hardware Standard turns lab devices into discoverable agent tools. For operators, the critical work is routing authority, safety limits, and audit paths before code touches hardware.

Anthropic Inference Hooks Put a Pre-Inference Gate at the Provider Layer: What It Means for Your Routing Architecture
Anthropic's new Inference Hooks let enterprise organizations intercept every governed Claude prompt before the model runs. For teams already filtering at the gateway layer, this creates a dual-gate architecture that changes where enforcement belongs.