← All articles

LLM API Production Hardening Checklist: Timeouts, Retries, Circuit Breakers, and Graceful Degradation

A unified checklist for hardening LLM API integrations in production: timeout budgets, retry policies with jitter, circuit breaker patterns, graceful degradation chains, rate limit absorption, streaming reliability, and health-check-driven routing.

· TheRouter

A 30-second answer: production LLM API hardening means layering five defense mechanisms, each handling a different failure mode. Timeouts prevent hung connections from blocking your application. Retries with jitter recover from transient errors without creating retry storms. Circuit breakers stop traffic to failing providers before cascading failures bring down your entire system. Graceful degradation chains route to cheaper or cached alternatives when primary models are unavailable. And health-check-driven routing makes all of this proactive rather than reactive. This checklist unifies what we have learned operating OpenAI-compatible routing across multiple providers.

OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.

Why a unified checklist matters

Individual reliability topics are well-documented. OpenAI publishes rate limit guidance. Anthropic documents error codes and SDK retry behavior. DeepSeek lists error codes with recommended actions. We have covered timeouts, retries, and idempotency and fallback routing separately.

What is missing is a single actionable checklist that connects all of these into a coherent production deployment. Teams end up with retries that fight their circuit breakers, timeout settings that defeat their fallback logic, and health checks that watch the wrong signals. This guide ties the pieces together.

The production hardening checklist

Before diving into each section, here is the checklist at a glance. Print it, paste it into your deployment runbook, and check each item before you ship.

#CategoryCheckPriority
1TimeoutsConnect timeout set (5-10s)Critical
2TimeoutsRead timeout set per use case (30-120s non-streaming)Critical
3TimeoutsTotal request timeout caps end-to-end latencyCritical
4TimeoutsStreaming per-chunk deadline configuredHigh
5RetriesOnly 429 and 5xx retriedCritical
6RetriesExponential backoff with full jitterCritical
7RetriesRetry-After header honoredCritical
8RetriesMax 3-5 attempts, retry budget <10% of trafficHigh
9RetriesRetries happen at one layer onlyHigh
10Circuit breakerFailure-rate threshold configured (20-30% over 60s)High
11Circuit breakerOpen state fails fast, no provider callsHigh
12Circuit breakerHalf-open probes test recoveryMedium
13DegradationModel downgrade chain definedHigh
14DegradationCached/queued fallback for total outageMedium
15Rate limitsDual RPM + TPM tracking at application layerHigh
16Rate limitsPre-request token estimationMedium
17StreamingPartial response handling and reconnectionMedium
18MonitoringLatency p50/p95/p99 per providerHigh
19MonitoringError rate and circuit breaker state alertsHigh
20MonitoringCost-per-request tracking and anomaly detectionMedium

Timeout strategy

Timeouts are the first line of defense. Without them, a single slow or unresponsive provider ties up connections, threads, and user attention indefinitely.

LLM APIs are slower than typical REST APIs. A generation request routinely takes 5-60 seconds depending on model size, input length, and output length. Standard HTTP client defaults (often 30 seconds total) will cause false timeouts on legitimate long-generation requests. But no timeout at all means a provider outage silently blocks your application.

Connect timeout guards the TCP/TLS handshake phase. Set this to 5-10 seconds. If a provider cannot accept a connection within this window, the endpoint is likely down or unreachable. No amount of waiting will fix a DNS resolution failure or a firewall drop.

Read timeout (or first-byte timeout for streaming) guards the period between sending the request and receiving the first response byte. For non-streaming requests, 30-120 seconds is reasonable depending on the model. For streaming, the first chunk should arrive within 10-30 seconds for most models; after that, set a per-chunk deadline of 15-30 seconds.

Total timeout caps the entire request lifecycle. This is your SLA backstop. If a request has not completed within your total budget (typically 60-180 seconds for non-streaming, higher for very long generations), terminate it regardless of partial progress.

import httpx

# Production timeout configuration for LLM API calls
client = httpx.Client(
    timeout=httpx.Timeout(
        connect=5.0,       # TCP/TLS handshake
        read=60.0,         # Wait for first byte / next chunk
        write=10.0,        # Sending the request body
        pool=10.0,         # Waiting for a connection from the pool
    )
)

Per-provider recommendations based on documented behavior:

ProviderConnectRead (non-streaming)Read (streaming first chunk)Notes
OpenAI5-10s60-120s15-30sLong-context models (GPT-5.5) may need higher read timeouts
Anthropic5-10s60-120s15-30sExtended thinking requests can take longer; 504 returned on server-side timeout
DeepSeek5-10s60-90s10-20sV4 Flash is typically fast; V4 Pro reasoning may need more time
DashScope5-10s60-90s10-20sOpenAI-compatible endpoint; same timeout logic applies

Retry policy with exponential backoff and jitter

Retries recover from transient failures. But naive retries cause more damage than the original failure. The canonical pattern for LLM APIs:

What to retry:

  • 429 (rate limited) with Retry-After header honored
  • 500 (internal server error) with backoff
  • 503 (service overloaded) with backoff
  • Connection reset, DNS timeout, TLS handshake failure

What to never retry:

  • 400 (bad request) will fail identically every time
  • 401 (authentication error) requires fixing the API key
  • 402 (billing/balance) requires adding funds
  • 403 (permission error) requires changing configuration
  • 413 (request too large) requires shrinking the request

Exponential backoff with full jitter prevents retry storms. Pure exponential backoff without jitter synchronizes all clients to retry at the same instant, recreating the thundering herd on every attempt:

import random
import time

def retry_with_jitter(func, max_attempts=3, base_delay=1.0, max_delay=32.0):
    for attempt in range(max_attempts):
        try:
            return func()
        except RetryableError as e:
            if attempt == max_attempts - 1:
                raise

            # Honor Retry-After header if present
            retry_after = getattr(e, 'retry_after', None)
            if retry_after:
                delay = float(retry_after)
            else:
                # Full jitter: random between 0 and exponential cap
                exp_delay = min(max_delay, base_delay * (2 ** attempt))
                delay = random.uniform(0, exp_delay)

            time.sleep(delay)

Retry budget prevents one degraded endpoint from consuming all your capacity. Set a global constraint: total retries should not exceed 10% of total requests at any time. If retry rate exceeds the budget, fail fast and route to fallback instead of continuing to hammer the degraded provider.

Single-layer retries prevent multiplicative explosion. If your application calls a service that calls another service, retries at every hop multiply. Three retries at each layer of a five-service chain produce 3^5 = 243 backend calls for a single user request. Pick one layer for retries, usually the outermost application layer or the routing gateway.

When using a routing layer like TheRouter, configure retries at the router level rather than in application code. The router sees all provider endpoints and can make smarter decisions about when to retry versus when to fail over to an alternative route. This avoids the double-retry problem where both your application and the gateway retry the same failed request.

Circuit breaker pattern for LLM providers

Retries handle transient failures. Circuit breakers handle systemic failures. The difference matters.

A circuit breaker monitors the failure rate over a rolling window and has three states:

Closed (normal operation): all requests pass through to the provider. The breaker tracks success/failure counts over the rolling window.

Open (tripped): the failure rate exceeded the threshold. All requests fail immediately without touching the provider. This gives the provider time to recover and prevents your application from wasting resources on requests that will fail.

Half-open (probing): after the cooldown period expires, the breaker allows a small number of test requests through. If they succeed, the breaker closes and normal traffic resumes. If they fail, the breaker reopens for another cooldown cycle.

For LLM APIs, configure these thresholds:

ParameterRecommended valueRationale
Rolling window60 secondsLong enough to detect sustained issues, short enough to react quickly
Failure threshold20-30% of requestsHigher than typical microservices because LLM APIs have baseline error rates
Consecutive failures to trip5-10Alternative trigger for low-traffic endpoints
Cooldown period30-60 secondsEnough time for provider recovery
Half-open probe count1-3 requestsMinimal traffic to test recovery
Reset on successAfter 2-3 consecutive half-open successesConfirm recovery is stable, not a single lucky request

LLM-specific circuit breaker triggers beyond standard HTTP error rates:

  • Latency degradation: trip when p95 latency exceeds 3x the baseline. A provider returning 200 OK but taking 90 seconds per request is effectively degraded.
  • Cost anomaly: trip when cost per request exceeds a configured threshold. This catches runaway agent loops and unexpected billing spikes.
  • Streaming interruption rate: trip when more than 30% of streaming responses disconnect before the final chunk.
// Conceptual circuit breaker for LLM provider routing
interface CircuitBreakerConfig {
  windowMs: number;          // Rolling window: 60000
  failureThreshold: number;  // Percentage: 0.25
  cooldownMs: number;        // Cooldown: 30000
  halfOpenProbes: number;    // Probes: 2
  latencyThresholdMs: number; // P95 cap: 45000
}

// When the breaker trips, route to the next provider in the
// fallback chain instead of returning an error to the user
function onCircuitOpen(provider: string, chain: string[]) {
  const next = chain.find(p => getCircuitState(p) === 'closed');
  if (next) {
    routeTo(next);  // Transparent fallback
  } else {
    serveDegraded(); // All providers degraded
  }
}

Graceful degradation chains

When circuit breakers trip, your application needs somewhere to go. Graceful degradation chains define the fallback path from your primary model through progressively cheaper or simpler alternatives, down to cached responses or user-visible degraded mode.

A practical degradation chain for a coding assistant:

LevelModelTriggerUser impact
L0 (primary)DeepSeek V4 ProNormal operationFull capability
L1 (fallback)DeepSeek V4 FlashL0 circuit open or latency >45sSlightly less accurate, much faster
L2 (budget)Qwen3.8 Flash via DashScopeL0 + L1 circuits openDifferent model, similar capability tier
L3 (cached)Cached responses for common queriesAll live providers degradedStale but available
L4 (degraded)Error message with ETAAll fallbacks exhaustedHonest about unavailability

The key design decisions:

Cross-provider fallback is more resilient than same-provider fallback. If OpenAI is down, falling back to another OpenAI model may not help. Falling back to Anthropic or DeepSeek via TheRouter's model fallback routing routes around provider-level outages.

Model downgrade is often better than provider outage. A user getting a response from a smaller model in 2 seconds is better served than waiting 30 seconds for a timeout from the large model. Configure downgrade thresholds based on latency, not just errors.

Cached responses need a freshness policy. For FAQ-style queries, a 1-hour cache is often acceptable. For real-time data questions, caching is not appropriate. Tag your queries with cacheability metadata.

Honest degradation builds trust. When all fallbacks fail, tell the user what happened and when you expect recovery. A clear "Our AI service is temporarily degraded; we expect recovery within 15 minutes" is better than a spinner that never resolves.

Rate limit handling and absorption

Rate limits are not errors; they are flow control. Production systems should absorb rate limits gracefully rather than treating them as failures.

Dual-axis tracking: LLM providers rate-limit on both RPM (requests per minute) and TPM (tokens per minute) simultaneously. You can stay within RPM while exceeding TPM, especially with long-context requests. Track both at your application layer, not just at the provider edge.

Pre-request token estimation prevents surprise TPM overruns. Use a tokenizer (tiktoken for OpenAI-compatible APIs, provider-specific libraries elsewhere) to estimate token count before sending the request. If the estimated tokens would exceed your remaining TPM budget, queue or shed the request rather than sending it and getting a 429.

Always set max_tokens to cap output length. Without this, a model generating an unusually long response can silently exhaust your TPM budget on a single request.

Burst absorption with queuing: instead of shedding requests that exceed your rate limit, queue them with a bounded wait time. A Redis-backed or in-memory priority queue with a 5-10 second maximum wait smooths burst traffic without losing requests:

import asyncio
from collections import deque

class TokenBucketLimiter:
    def __init__(self, rpm: int, tpm: int):
        self.rpm = rpm
        self.tpm = tpm
        self.request_tokens = rpm
        self.token_tokens = tpm
        self.queue: deque = deque()

    async def acquire(self, estimated_tokens: int, timeout: float = 5.0):
        """Wait up to timeout seconds for capacity."""
        deadline = asyncio.get_event_loop().time() + timeout
        while asyncio.get_event_loop().time() < deadline:
            if self.request_tokens > 0 and self.token_tokens >= estimated_tokens:
                self.request_tokens -= 1
                self.token_tokens -= estimated_tokens
                return True
            await asyncio.sleep(0.1)
        return False  # Shed or route to fallback

Streaming reliability

Streaming (SSE) responses add reliability challenges that batch requests do not have:

Partial response handling: a streaming response can disconnect midway. Your application should track how much output was received and whether it includes a completion marker (the [DONE] event in OpenAI-compatible APIs, or message_stop in Anthropic's format). A partial response without the stop marker needs to be flagged as incomplete.

Reconnection strategy: do not blindly reconnect and re-send the same prompt. The provider may have already processed part of the request and billed for it. Instead, mark the partial response as incomplete and either present it to the user with a warning or start a new request from scratch with adjusted context.

Per-chunk timeout: set a deadline for each SSE chunk, not just for the first byte. A provider that sends the first chunk in 2 seconds but then stalls for 60 seconds between chunks is effectively degraded. A 15-30 second per-chunk deadline catches this.

Error events mid-stream: both OpenAI and Anthropic can return error events after the initial 200 response. Your SSE parser must handle error event types that arrive mid-stream, not just HTTP-level error responses.

Health checks and proactive routing

Reactive resilience (retry after failure, trip after errors) is the baseline. Proactive resilience (route away from degraded providers before they affect users) is the next level.

Synthetic health probes: send lightweight test requests to each provider on a schedule (every 30-60 seconds). Use a small, fast model and a short prompt. Track response time and success rate. If a provider's probe latency exceeds 2x baseline or probes start failing, reduce its routing weight before user traffic is affected.

Provider status pages: monitor official status pages (status.openai.com, status.claude.com, status.deepseek.com) programmatically. When a provider reports a degraded or major outage, pre-emptively route traffic away.

Response header monitoring: OpenAI returns x-ratelimit-remaining-requests and x-ratelimit-remaining-tokens headers on every response. Use these to predict when you will hit limits and proactively throttle or reroute before getting a 429.

When using TheRouter as your routing layer, health-check routing is built into the model fallback configuration. The router monitors provider health across all traffic passing through it and routes around degraded endpoints without requiring application-level health check code.

Monitoring and SLO definition

You cannot harden what you do not measure. Define SLOs for your LLM API integration and alert on breaches.

Key metrics to track:

MetricMeasurementSLO example
Request success rate1 - (error_count / total_count)>99.5% over 5-minute windows
Latency p50Median end-to-end request time<5s for flash models, <15s for flagship
Latency p9595th percentile request time<15s for flash, <45s for flagship
Latency p9999th percentile request time<30s for flash, <90s for flagship
Circuit breaker open timeTotal seconds breakers are open per hour<300s/hour
Cost per requestTotal spend / total requestsBelow budget threshold
Retry rateretry_count / total_count<5% sustained, <10% peak
Streaming completion ratecomplete_streams / started_streams>99%

Alert on:

  • Success rate drops below SLO for 2 consecutive windows
  • Any circuit breaker transitions to open
  • Cost per request exceeds 2x the 7-day rolling average
  • Retry rate exceeds 10% for more than 5 minutes

Putting it all together

The defense layers interact. Here is how they compose:

  1. Request arrives at your application
  2. Pre-request check: token estimation, rate limit budget check. If over budget, queue or shed.
  3. Route selection: health-check-driven routing picks the best available provider/model
  4. Timeout enforcement: connect + read + total timeouts wrap the API call
  5. Failure handling: if the request fails, classify the error
  6. Retry decision: retryable errors go through exponential backoff with jitter, within the retry budget
  7. Circuit breaker check: if the provider's circuit is open, skip to fallback immediately
  8. Fallback routing: if retries are exhausted or the circuit is open, route to the next model in the degradation chain
  9. Degraded mode: if all providers are down, serve cached responses or return an honest error
  10. Telemetry: log every decision point for monitoring and SLO tracking

This is not a recommendation to build all of this from scratch. An LLM routing layer handles steps 3, 4, 6, 7, and 8 at the infrastructure level, letting your application code focus on steps 1, 2, 5, 9, and 10.

FAQ

What timeout should I set for extended thinking / reasoning models?

Reasoning models like DeepSeek R1 or Claude with extended thinking can take 60-180 seconds to produce output. Set read timeouts to at least 120s for these models. Use streaming to get partial visibility into progress. If you use a total timeout, set it to 3x the expected generation time for the model and prompt length.

Should I retry tool calls and function calls?

Only if the tool execution sink is idempotent. If a tool call sends an email, creates a ticket, or writes to a database, retrying the request can duplicate the side effect. Check whether the previous call executed before retrying. Use idempotency keys when the provider supports them.

How do I handle rate limits across multiple API keys?

Track RPM and TPM per key, not per application. If you distribute traffic across multiple keys (for higher aggregate throughput), each key has its own limits. Your rate limiter needs to track each key independently and route to keys with remaining capacity.

What is the difference between application-level and gateway-level retries?

Application-level retries happen in your code. Gateway-level retries happen in the routing layer (like TheRouter or a reverse proxy). Use one or the other, not both. If both retry, you get multiplicative retry storms. Gateway-level retries are generally preferable because the gateway sees all traffic and can make smarter decisions about provider health.

How often should circuit breaker thresholds be tuned?

Review thresholds monthly. Provider reliability characteristics change over time. A threshold that was appropriate when a provider had 99.5% uptime may be too sensitive if the provider improves to 99.9%, or too loose if it degrades. Use your monitoring data to tune: if the breaker trips more than once a day on false positives, raise the threshold. If users report degraded experiences before the breaker trips, lower it.

When should I use a routing layer versus building retry and fallback logic in my application?

If you call one provider with one model, application-level retry logic is fine. If you call multiple providers, use multiple models, or need circuit breakers and health-check routing, a dedicated routing layer removes significant complexity from your application code and centralizes the resilience logic where it can be monitored and tuned independently.


Sources cited in this guide:

Help & contact