Anthropic Unifies Claude API Rate Limits Across Tiers: What Sonnet/Haiku/Opus Parity Means for Routing Teams

Anthropic's June 26 rate limit overhaul gives Sonnet and Haiku the same RPM, ITPM, and OTPM as Opus at every tier — and consolidates five usage tiers into three. Here's what the new Start/Build/Scale structure means for fallback routing policy and spend cap management.

TheRouter Newsroomvia Anthropic Platform Docs
Anthropic Claude API rate limits tier unification routing — Start Build Scale usage tier consolidation editorial

On June 26, Anthropic quietly shipped the most consequential rate limit update since the Mythos 5 rollout: Claude Sonnet 4.x and Claude Haiku 4.5 now carry identical RPM, ITPM, and OTPM ceilings to Claude Opus 4.x at every usage tier. Alongside that, five legacy tiers were merged into three — Start, Build, and Scale — each with a hard monthly spend cap. If your routing policy was built around separate per-model buckets, those assumptions no longer hold.

What changed on June 26

Before this update, Haiku had a separate, lower token budget from Sonnet and Opus. Teams building fallback chains often routed high-volume, lower-stakes traffic to Haiku specifically to preserve Opus headroom inside a shared tier. That arbitrage is gone.

The new limits are uniform across Opus 4.x, Sonnet 4.x, and Haiku 4.5 within each tier:

Usage tierModels (Opus 4.x / Sonnet 4.x / Haiku 4.5)RPMITPMOTPM
StartAll three1,0002,000,000400,000
BuildAll three5,0005,000,0001,000,000
ScaleAll three10,00010,000,0002,000,000

Note: rate limits are applied separately per model — Opus traffic doesn't eat into Sonnet's bucket — but within each model the limits are now symmetric across tiers.

Claude Fable 5 remains on its own schedule (Start: 500K ITPM, Build: 1.5M, Scale: 4M), reflecting its higher compute cost.

The consolidation also removed intermediate tiers. Organizations that were on a mid-tier are automatically assigned to the nearest tier; Anthropic confirmed no organization received a lower limit than before.

Spend caps per tier

The three-tier structure comes with monthly spend ceilings:

TierMonthly spend cap
Start$500
Build$1,000
Scale$200,000

Custom tier organizations (contract-negotiated) have no cap. For Scale-tier teams, $200K/month is substantial headroom — but teams pushing into the Custom tier through a single API key cluster should note the accelerator-limit note in the docs: sudden traffic spikes trigger 429s independent of the tier limit, so ramp gradually.

The cache-aware ITPM advantage

The most underappreciated part of Anthropic's rate limit design is what doesn't count. For all current Claude models except the retired Haiku 3.5, cache_read_input_tokens are excluded from ITPM accounting. Only uncached input tokens (input_tokens + cache_creation_input_tokens) consume your token budget.

Practically: at the Build tier with 5M ITPM and an 80% cache hit rate, your effective throughput is ~25M total input tokens per minute. This is a permanent structural advantage for routing workloads that share large system prompts, tool definitions, or document contexts across many requests.

Routing teams using a shared-cache architecture — where a common system prompt is written once and then read from cache on subsequent requests — can sustain significantly higher concurrency than the nominal limit implies. The Rate Limits API (/manage-claude/rate-limits-api) lets you poll remaining headroom programmatically, making it feasible to build dynamic routing that steps down to a lighter model only when actual uncached ITPM pressure is real, not predicted.

What this means for fallback routing policy

Under the old multi-tier, per-model regime, many teams structured fallbacks like this:

  1. Primary: Opus 4.x (best quality, lower quota ceiling)
  2. Fallback: Sonnet 4.x (mid tier, separate bucket)
  3. Emergency fallback: Haiku 4.5 (high-RPM cheapest, isolated limit)

This preserved Opus headroom by routing volume traffic to Haiku's separate and higher-limit bucket.

With parity, the model-segregation strategy no longer moves the budget needle. Each model still has independent buckets, so Haiku traffic at 10K RPM doesn't block Opus traffic at 10K RPM — limits remain per-model, not shared. But the relative "save Haiku quota for overflow" trick is obsolete, because Haiku now has the same ceiling as Opus at each tier.

The correct routing logic now centers on:

  • Cost efficiency per task: route Haiku for high-volume summarization; Opus/Fable for reasoning chains. The budget is symmetric, so the decision is purely cost and quality.
  • Cache hit rate: design system prompts and tool schemas for cache reuse to expand effective ITPM without tier upgrades.
  • Spend cap proximity: monitor monthly spend against tier cap; a hard pause at $500/$1K/$200K is a reliability risk — set a workspace-level sub-limit below the cap to avoid surprise outages.

The AWS note operators often miss

Organizations accessing Claude through Claude Platform on AWS (Bedrock-path) are placed on the Start tier and do not move between tiers automatically. To request higher limits, you must contact your Anthropic account representative. Per-workspace rate limit configuration and fast mode are also unavailable on that path.

If you're routing Claude traffic through an AWS-native setup and expecting automatic tier progression, verify with your account team.

What to check now

  1. Review your workspace limits in the Claude Console — confirm which tier you're on and whether the new parity changes your fallback assumptions.
  2. Audit your cache hit rate on the Usage page. If it's below 50% for workloads with shared system prompts, you're leaving effective ITPM on the table.
  3. Set a workspace-level spend sub-limit below your tier cap to avoid hard pauses at the cap ceiling.
  4. Check your fallback chain logic: if model-diversity was the reason for your Haiku fallback, update the rationale to cost/quality rather than budget segregation.
  5. AWS operators: confirm tier assignment and whether you need to file for a limit increase through your account rep.

The June 26 change is a structural simplification — fewer tiers, symmetric limits, and a clean spend cap. For routing teams, the implication is that model selection should now be driven by cost and capability rather than by which model had the biggest limit bucket.

Help & contact