Anthropic Doubles Claude Code Limits, Adds SpaceX GPUs: Routing Impact

Anthropic doubled Claude Code rate limits and raised Opus API ceilings with 220K GPUs at SpaceX Colossus 1. Learn how this changes fallback triggers, scheduling, and multi-provider routing strategy.

Published via Anthropic

Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

Abstract diagram of AI compute infrastructure capacity expansion with routing policy decision layers

Rate limits are not just pricing friction — they are the architectural boundary that determines whether your primary provider can absorb a traffic spike or hands the request off to a fallback. When Anthropic doubled Claude Code's limits on May 6, 2026 and simultaneously disclosed 220,000+ new NVIDIA GPUs coming online at SpaceX's Colossus 1 data center, the operational question for engineering teams was not "interesting, more capacity." It was: does this change how I configure my routing policy?

The short answer is yes, in three specific ways.

What changed

Anthropic announced three concrete changes, effective May 6, 2026:

1. Claude Code five-hour rate limits doubled. For Pro, Max, Team, and seat-based Enterprise plans, the rolling five-hour usage allowance for Claude Code is now 2× what it was. This is not a temporary promotion; it reflects the compute capacity coming online.

2. Peak-hours throttling removed for Pro and Max. Previously, Claude Code users on Pro and Max plans experienced reduced limits during peak demand periods. That reduction is gone. Limits are now flat across all hours, which matters for teams with workloads that run in overlap with US business hours.

3. Claude Opus API rate limits raised considerably. The announcement included an updated rate-limit table for Claude Opus models. Opus is typically the highest-cost, highest-capability tier — the one most teams use for complex reasoning, long-context tasks, and agent orchestration. Raising its API rate limits directly affects how reliably teams can run Opus at scale without needing to fall back to Sonnet or flash-tier models.

The compute backstory: Anthropic signed an agreement to use all compute capacity at SpaceX's Colossus 1 data center in Memphis — over 300 MW, more than 220,000 NVIDIA GPUs, coming online within a month of the announcement. This joins AWS Trainium, Google TPUs, a $30B Azure/NVIDIA strategic partnership, and a planned 5 GW deal with Amazon and Google. Anthropic's compute stack is no longer single-provider; it is diversified across five major compute sources.

Why it matters for AI engineering teams

Burst ceiling changed. The most immediate operational impact is that the effective burst ceiling for Claude Code and Opus API is higher. Teams that previously hit limits during long agent runs, batch processing windows, or overnight automation pipelines should see those constraints ease. This does not mean limits are gone — it means the headroom before you hit them is larger.

Fallback trigger rate shifts. If you configured fallback routing to kick in when Anthropic returns a 429 (rate limit exceeded), you will see fewer triggers. For teams that built fallback chains specifically to handle Anthropic capacity constraints (e.g., falling back to Gemini or DeepSeek on Anthropic 429s), the practical fallback rate should decrease. This is worth measuring in your observability stack — if you were using fallback frequency as a proxy for provider health, a reduced rate is now baseline-normal, not a sign of unusually low demand.

Peak-hours removal changes your scheduling assumptions. Teams that deliberately scheduled non-urgent AI workloads for off-peak hours (nights, weekends) to avoid Pro/Max throttling can now stop doing that. If you have cron jobs, batch runs, or CI pipelines scheduled at non-peak times specifically to avoid Anthropic rate limits, you can reconsider that scheduling constraint.

Multi-compute stack is a reliability signal. The diversification of Anthropic's compute across AWS, Google, SpaceX, and eventually Azure/NVIDIA means that a single data center failure is less likely to cause a Anthropic-wide outage. For teams that include Anthropic in multi-provider routing for reliability (not just cost), this is a structural improvement in their provider's resilience posture.

The router/operator angle

Three routing policy decisions this compute expansion creates:

1. Recalibrate your Opus fallback threshold. If you configured a hard fallback from Claude Opus to a cheaper model at a specific token-count-per-hour threshold, that threshold was likely set based on old rate limits. Review whether the threshold still makes sense, or whether you can safely raise it to take advantage of the expanded headroom. Over-aggressive fallback from Opus to Sonnet costs you quality; under-aggressive fallback costs you money. Rebalance with current data.

2. Audit whether peak-hour scheduling is still necessary. Any pipeline that artificially avoids Anthropic peak hours to sidestep throttling should be audited. If you can simplify your scheduling logic (removing off-peak constraints), that is one less moving part in your infrastructure.

3. Consider the multi-provider compute diversification when scoring Anthropic reliability. If your provider-selection logic includes a reliability score or provider health weight, the compute diversification across five major sources is worth factoring in. Anthropic's reliability posture has improved structurally, not just in terms of raw capacity.

Routing tier framework — post May 6 limits:

ScenarioRecommended approach
Opus hitting rate limits in agentic pipelineRaise fallback threshold; measure actual 429 rate before config change
CI/cron scheduled off-peak to avoid throttleRelax scheduling constraint; simplify pipeline
Fallback chain: Opus → Sonnet → DeepSeekReview fallback trigger frequency; confirm new limits in observability
Evaluating Anthropic provider reliabilityFactor in multi-compute-source architecture (SpaceX + AWS + Google + Azure)

One latency note: More compute capacity does not by itself improve inference latency on individual requests. Time-to-first-token and throughput per request depend on infrastructure beyond raw GPU count. If your routing policy includes latency-based provider selection, the May 6 changes affect capacity and rate limits, not necessarily per-request latency characteristics.

What TheRouter users should watch or try

If you route Claude Opus through TheRouter, the May 6 rate limit increases mean your provider-level rate limit configuration may now be more conservative than necessary. If you set a requests-per-minute or tokens-per-minute ceiling when Anthropic's limits were lower, review whether that ceiling still reflects your actual provider headroom.

Observability check: pull your Anthropic 429 error rate over the past 30 days and compare it against your pre-May-6 baseline. If the fallback trigger rate dropped, the new limits are working. If it did not, there may be account-tier or plan mismatches worth investigating with Anthropic support.

The broader pattern here is important for multi-provider routing strategy: compute capacity expansions at major providers create asymmetries in provider reliability that do not show up in benchmark comparisons or pricing tables. They show up in production 429 rates, burst availability, and the accuracy of your fallback frequency assumptions. This is exactly the kind of provider-level operational signal that routing infrastructure should track — and act on when it changes.

Help & contact