← All articles

LLM API Provider Reliability and Uptime Comparison 2026 H2: Outage Patterns, SLA Gaps, and Fallback Architecture

We compared uptime, outage frequency, and incident resolution times across OpenAI, Anthropic, Google, DeepSeek, DashScope, and SiliconFlow APIs through the first eight months of 2026. AI APIs remain the least reliable API category tracked by independent monitors. Here is what the data shows and how multi-provider fallback routing changes the math.

· TheRouter

Every provider goes down. The question is whether your application goes down with it.

We spent the first eight months of 2026 watching LLM API providers ship features at a pace that makes SaaS release cycles look glacial, and break production in the process. This post collects the reliability data we have tracked across OpenAI, Anthropic, Google, DeepSeek, DashScope, and SiliconFlow and turns it into a decision framework for fallback routing.

OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.

The Reliability Landscape in 2026

AI and ML APIs are the least reliable API category tracked across 215+ services, according to the Nordic APIs Reliability Report covering October 2025 through February 2026. Stripe runs at roughly 99.99% uptime. Linear posts 99.96%. The best LLM API providers hover around 99.85–99.90%, and the worst dip below 98%.

The ModelUptime May 2026 report — an independent cross-provider monitor tracking 9 providers and 29 models — measured 177 incidents in a single month with an average uptime of 99.28%.

These are not numbers from an industry that has reliability figured out.

Provider-by-Provider Reliability Profiles

The following table aggregates data from public status pages, independent monitors (ModelUptime, isDown.app, StatusGator, IncidentHub), and the Nordic APIs report.

ProviderSelf-Reported UptimeIndependent Uptime (May 2026)Incidents (May 2026)Severe IncidentsStatus Page
OpenAI99.93% (Jun–Sep 2026)99.86%61status.openai.com
Anthropic99.01–99.59% (Jul–Aug 2026)99.85%181status.claude.com
GoogleNot published98.22%3512status.cloud.google.com
DeepSeek99.88% (Jun–Sep 2026)98.16%505status.deepseek.com
DashScopeNot publishedNot independently tracked——Alibaba Cloud status
SiliconFlowNot publishedNot independently tracked———

Sources: status.openai.com, status.claude.com/uptime, status.deepseek.com, ModelUptime May 2026, Nordic APIs Report

Key observations:

  • The self-reported vs. independent gap is real. DeepSeek claims 99.88% on its status page while ModelUptime measured 98.16% in the same period. Providers choose what makes it onto their status page. Independent monitors catch what they miss.
  • Incident count does not equal downtime. Anthropic had 18 incidents to OpenAI's 6, but their uptime numbers were nearly identical (99.85% vs. 99.86%). Anthropic resolves issues faster on average.
  • Google had the most severe incidents. 12 out of 35 incidents were classified as severe, the highest ratio of any provider. Gemini API reliability has been volatile.
  • DashScope and SiliconFlow lack independent monitoring. This is not necessarily bad — Alibaba Cloud's infrastructure is mature — but it means we cannot verify claims the same way we can for providers with public incident histories.

Outage Patterns and Incident Frequency

OpenAI

OpenAI logged 11 incidents in 28 days during January 2026 — one every 2.5 days. By mid-2026, incident frequency had improved significantly: 6 incidents in May, and their status page shows 99.93% uptime for the June–September window. Most incidents resolve in 30 to 90 minutes.

OpenAI offers uptime commitments only through its Scale Tier program. Standard API customers get no contractual SLA. As of July 2026, Scale Tier traffic automatically spills over to Fast mode during capacity constraints.

Anthropic

Anthropic's Claude API had a rough Q1 2026. A 30-hour resolution cycle for Claude Opus 4.5 in late January set the tone. The March 26–27 networking incident caused elevated error rates across Opus 4.6 and Sonnet 4.6. By mid-2026, monthly uptime improved from 99.01% (July) to 99.59% (August).

The most recent incident at time of writing was September 11, 2026: elevated errors for Claude Mythos 5.1 and Claude Fable 5.1, per isDown.app. Anthropic does not publish a formal API uptime SLA for standard customers.

Google

Google's Gemini API had the worst severe-incident ratio in May 2026: 12 severe incidents out of 35 total. Independent monitoring showed 98.22% uptime, well below the other major Western providers. Google does not publish a Gemini-specific uptime commitment; Vertex AI customers can reference Google Cloud's general SLAs.

DeepSeek

DeepSeek was the weakest provider in ModelUptime's May 2026 report at 98.16% uptime with 50 incidents. On August 4, 2026, two degraded-performance incidents hit the V4 Flash API within the same day, the first lasting 1 hour 18 minutes. DeepSeek's own status page reports 99.88% for the June–September window — a significant gap from independent measurements.

DeepSeek does not publish rate-limit or uptime SLAs. During peak demand periods (particularly after major model launches), the API has historically experienced extended capacity constraints.

DashScope (Alibaba Cloud Model Studio)

DashScope benefits from Alibaba Cloud's production-grade infrastructure. It does not maintain a public incident history specific to Model Studio, and no major independent monitor tracks DashScope API uptime separately from Alibaba Cloud's broader service health.

In our experience routing traffic through DashScope, we have observed stable performance during normal operating conditions, with occasional latency spikes during Chinese business hours (09:00–18:00 CST) that correlate with high regional demand.

SiliconFlow

SiliconFlow does not publish a public status page or incident history. As a younger platform focused on cost-efficient inference, their reliability track record is harder to evaluate independently. We have observed generally stable API responses for popular models, with occasional 429 (rate limit) responses on free-tier models during peak hours.

SLA Reality Check

Most LLM API providers do not offer contractual uptime SLAs to standard API customers. Here is what actually exists:

ProviderPublished SLAApplies ToCompensation
OpenAIYes (Scale Tier only)Scale Tier customersCredits for SLA violations
AnthropicNo public SLA——
GoogleVertex AI SLA onlyGoogle Cloud customersCloud credit per SLA terms
DeepSeekNo public SLA——
DashScopeAlibaba Cloud SLAAlibaba Cloud customersPer cloud agreement
SiliconFlowNo public SLA——

The gap between "we post a status page" and "we will compensate you when it breaks" is enormous. For most providers, a status page is a courtesy, not a commitment.

Three Failure Modes to Plan For

1. Hard downtime

The API returns 5xx errors or times out. This is the easy one to detect — your monitoring catches it immediately. The problem is that most teams have no automated fallback. Engineers open the status page, wait, and context-switch.

For a 50-person engineering team at $80–$150/hour, each hour of a critical dependency outage costs $4,000–$7,500 in lost productivity, per analysis by BuildMVPFast.

2. Degraded quality

The API returns 200 OK, but output is wrong, slow, or truncated. Health checks pass. The model responds — just badly. Latency spikes from 500ms to 8 seconds. Structured responses start failing schema validation. This failure mode is harder to catch and more insidious than a clean outage.

3. Cascade failures

Your AI provider is fine, but something upstream broke. The October 2025 AWS DynamoDB incident cascaded into 141 affected services. Cloudflare's November outage brought down parts of OpenAI itself. Your dependency chain is deeper than you think.

Fallback Architecture: The Multi-Provider Routing Solution

The data makes the case for multi-provider routing better than any marketing pitch could. When your primary provider goes down, your application should automatically route to a secondary provider — not page an engineer.

Here is a minimal fallback configuration using an OpenAI-compatible routing layer:

from openai import OpenAI

# Primary: route through TheRouter with fallback configured
client = OpenAI(
    base_url="https://api.therouter.ai/v1",
    api_key="your-therouter-key"
)

response = client.chat.completions.create(
    model="openai/gpt-4o",  # primary provider
    messages=[{"role": "user", "content": "Hello"}],
    # TheRouter handles fallback to anthropic/claude-sonnet-4.6
    # if OpenAI returns 5xx or exceeds latency threshold
)

With provider fallback routing, the application code never changes. The routing layer absorbs provider instability and returns a response from whichever provider is healthy.

Recommended Fallback Chains

Based on the reliability data above, here are fallback chains that maximize coverage:

Primary Use CasePrimaryFallback 1Fallback 2Rationale
General chatopenai/gpt-4oanthropic/claude-sonnet-4.6deepseek/deepseek-v4-flashDiverse infrastructure, different failure domains
Coding tasksanthropic/claude-opus-4.6deepseek/deepseek-v4-proopenai/gpt-4oOpus has highest coding benchmarks; DeepSeek is cost-effective backup
High-throughputdeepseek/deepseek-v4-flashdashscope/qwen3.7-plussiliconflow/deepseek-v4-flashFlash models for cost; DashScope on different infrastructure
Chinese marketdashscope/qwen3.8-maxdeepseek/deepseek-v4-prosiliconflow/qwen3.7-maxAll support Chinese well; different hosting infrastructure

The key principle: never stack fallbacks on the same infrastructure. OpenAI and Microsoft models share Azure infrastructure. DashScope-hosted DeepSeek and direct DeepSeek API share DeepSeek's backend. Effective fallback chains cross infrastructure boundaries.

Decision Framework: Minimum Fallback Depth by Reliability Requirement

Availability TargetRequired Fallback DepthConfiguration
99% (7.3 hours/month downtime OK)No fallback neededSingle provider, accept occasional outages
99.9% (43 minutes/month)1 fallback providerPrimary + 1 backup on different infrastructure
99.95% (22 minutes/month)2 fallback providersPrimary + 2 backups, health-check routing
99.99% (4 minutes/month)2+ fallbacks + active health monitoringMulti-provider with automated failover and latency-based routing

The math: if Provider A has 99.85% uptime and Provider B has 99.85% uptime independently, routing across both gives you approximately 99.9998% uptime — assuming their failures are uncorrelated. Correlated failures (shared infrastructure, shared upstream dependencies) reduce this benefit.

Building Your Observability Stack

Detecting degradation before your users do requires AI-specific monitoring:

  1. Track per-provider latency distributions, not just averages. A provider with stable p50 but volatile p95 is degrading intermittently.
  2. Monitor error rates by error code. A spike in 429s means rate limiting; a spike in 500s means infrastructure failure. The response is different.
  3. Validate output quality. A 200 OK with garbage output is worse than a clean 503. Schema validation on structured responses catches this.
  4. Set up status page webhooks. Every provider listed above publishes status page updates. Subscribe to them. A 10-minute heads-up on a developing incident beats discovering it through user complaints.

For a detailed comparison of LLM monitoring tools, see our LLM API observability and monitoring tools comparison.

FAQ

Which LLM API provider is the most reliable?

No single provider holds the top spot consistently. Microsoft led ModelUptime's May 2026 rankings at 99.90%, but Microsoft only had one model tracked (Phi-4). Among major providers with broad model coverage, OpenAI (99.86%) and Anthropic (99.85%) were within 0.01% of each other. Month-to-month variation is significant.

Should I switch providers based on reliability data?

Switching providers based on a single month of data is premature. A better approach is multi-provider routing: keep your primary provider for quality reasons and add a fallback for reliability. See our LLM API fallback routing guide for implementation details.

How do I monitor my LLM API provider's uptime?

Subscribe to your provider's status page (all major providers have one). Layer independent monitoring on top: ModelUptime, isDown.app, and StatusGator all track LLM API providers. For your own traffic, instrument per-request latency and error rates at the application level.

Why is there a gap between self-reported and independently measured uptime?

Providers decide what qualifies as an "incident" on their status page. A 10-minute latency spike might not appear on the status page but will show up in independent monitors. Self-reported uptime also excludes "degraded performance" from the uptime calculation more aggressively than independent monitors do.

Does multi-provider routing increase latency?

A well-configured routing layer adds 1–5ms of overhead per request for provider selection. This is negligible compared to the 500–3000ms typical LLM response time. The latency cost of routing is far smaller than the latency cost of waiting for a degraded provider to respond slowly.

How do I choose fallback providers for my use case?

Prioritize infrastructure diversity over model similarity. Your fallback should run on different infrastructure than your primary. OpenAI and Azure OpenAI share infrastructure. Direct DeepSeek API and DashScope-hosted DeepSeek share DeepSeek's backend. Cross those boundaries. See the error handling reference for details on how different providers fail.


Reliability data in this post is based on publicly available status pages and independent monitoring reports. Uptime figures are approximations based on incident reports, not SLA-grade measurements. Provider reliability changes over time — check the linked status pages for current conditions.

Help & contact