LLM API Production Hardening Checklist: Timeouts, Retries, Circuit Breakers, and Graceful Degradation
A unified checklist for hardening LLM API integrations in production: timeout budgets, retry policies with jitter, circuit breaker patterns, graceful degradation chains, rate limit absorption, streaming reliability, and health-check-driven routing.
A 30-second answer: production LLM API hardening means layering five defense mechanisms, each handling a different failure mode. Timeouts prevent hung connections from blocking your application. Retries with jitter recover from transient errors without creating retry storms. Circuit breakers stop traffic to failing providers before cascading failures bring down your entire system. Graceful degradation chains route to cheaper or cached alternatives when primary models are unavailable. And health-check-driven routing makes all of this proactive rather than reactive. This checklist unifies what we have learned operating OpenAI-compatible routing across multiple providers.
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
Why a unified checklist matters
Individual reliability topics are well-documented. OpenAI publishes rate limit guidance. Anthropic documents error codes and SDK retry behavior. DeepSeek lists error codes with recommended actions. We have covered timeouts, retries, and idempotency and fallback routing separately.
What is missing is a single actionable checklist that connects all of these into a coherent production deployment. Teams end up with retries that fight their circuit breakers, timeout settings that defeat their fallback logic, and health checks that watch the wrong signals. This guide ties the pieces together.
The production hardening checklist
Before diving into each section, here is the checklist at a glance. Print it, paste it into your deployment runbook, and check each item before you ship.
| # | Category | Check | Priority |
|---|---|---|---|
| 1 | Timeouts | Connect timeout set (5-10s) | Critical |
| 2 | Timeouts | Read timeout set per use case (30-120s non-streaming) | Critical |
| 3 | Timeouts | Total request timeout caps end-to-end latency | Critical |
| 4 | Timeouts | Streaming per-chunk deadline configured | High |
| 5 | Retries | Only 429 and 5xx retried | Critical |
| 6 | Retries | Exponential backoff with full jitter | Critical |
| 7 | Retries | Retry-After header honored | Critical |
| 8 | Retries | Max 3-5 attempts, retry budget <10% of traffic | High |
| 9 | Retries | Retries happen at one layer only | High |
| 10 | Circuit breaker | Failure-rate threshold configured (20-30% over 60s) | High |
| 11 | Circuit breaker | Open state fails fast, no provider calls | High |
| 12 | Circuit breaker | Half-open probes test recovery | Medium |
| 13 | Degradation | Model downgrade chain defined | High |
| 14 | Degradation | Cached/queued fallback for total outage | Medium |
| 15 | Rate limits | Dual RPM + TPM tracking at application layer | High |
| 16 | Rate limits | Pre-request token estimation | Medium |
| 17 | Streaming | Partial response handling and reconnection | Medium |
| 18 | Monitoring | Latency p50/p95/p99 per provider | High |
| 19 | Monitoring | Error rate and circuit breaker state alerts | High |
| 20 | Monitoring | Cost-per-request tracking and anomaly detection | Medium |
Timeout strategy
Timeouts are the first line of defense. Without them, a single slow or unresponsive provider ties up connections, threads, and user attention indefinitely.
LLM APIs are slower than typical REST APIs. A generation request routinely takes 5-60 seconds depending on model size, input length, and output length. Standard HTTP client defaults (often 30 seconds total) will cause false timeouts on legitimate long-generation requests. But no timeout at all means a provider outage silently blocks your application.
Connect timeout guards the TCP/TLS handshake phase. Set this to 5-10 seconds. If a provider cannot accept a connection within this window, the endpoint is likely down or unreachable. No amount of waiting will fix a DNS resolution failure or a firewall drop.
Read timeout (or first-byte timeout for streaming) guards the period between sending the request and receiving the first response byte. For non-streaming requests, 30-120 seconds is reasonable depending on the model. For streaming, the first chunk should arrive within 10-30 seconds for most models; after that, set a per-chunk deadline of 15-30 seconds.
Total timeout caps the entire request lifecycle. This is your SLA backstop. If a request has not completed within your total budget (typically 60-180 seconds for non-streaming, higher for very long generations), terminate it regardless of partial progress.
import httpx
# Production timeout configuration for LLM API calls
client = httpx.Client(
timeout=httpx.Timeout(
connect=5.0, # TCP/TLS handshake
read=60.0, # Wait for first byte / next chunk
write=10.0, # Sending the request body
pool=10.0, # Waiting for a connection from the pool
)
)
Per-provider recommendations based on documented behavior:
| Provider | Connect | Read (non-streaming) | Read (streaming first chunk) | Notes |
|---|---|---|---|---|
| OpenAI | 5-10s | 60-120s | 15-30s | Long-context models (GPT-5.5) may need higher read timeouts |
| Anthropic | 5-10s | 60-120s | 15-30s | Extended thinking requests can take longer; 504 returned on server-side timeout |
| DeepSeek | 5-10s | 60-90s | 10-20s | V4 Flash is typically fast; V4 Pro reasoning may need more time |
| DashScope | 5-10s | 60-90s | 10-20s | OpenAI-compatible endpoint; same timeout logic applies |
Retry policy with exponential backoff and jitter
Retries recover from transient failures. But naive retries cause more damage than the original failure. The canonical pattern for LLM APIs:
What to retry:
- 429 (rate limited) with
Retry-Afterheader honored - 500 (internal server error) with backoff
- 503 (service overloaded) with backoff
- Connection reset, DNS timeout, TLS handshake failure
What to never retry:
- 400 (bad request) will fail identically every time
- 401 (authentication error) requires fixing the API key
- 402 (billing/balance) requires adding funds
- 403 (permission error) requires changing configuration
- 413 (request too large) requires shrinking the request
Exponential backoff with full jitter prevents retry storms. Pure exponential backoff without jitter synchronizes all clients to retry at the same instant, recreating the thundering herd on every attempt:
import random
import time
def retry_with_jitter(func, max_attempts=3, base_delay=1.0, max_delay=32.0):
for attempt in range(max_attempts):
try:
return func()
except RetryableError as e:
if attempt == max_attempts - 1:
raise
# Honor Retry-After header if present
retry_after = getattr(e, 'retry_after', None)
if retry_after:
delay = float(retry_after)
else:
# Full jitter: random between 0 and exponential cap
exp_delay = min(max_delay, base_delay * (2 ** attempt))
delay = random.uniform(0, exp_delay)
time.sleep(delay)
Retry budget prevents one degraded endpoint from consuming all your capacity. Set a global constraint: total retries should not exceed 10% of total requests at any time. If retry rate exceeds the budget, fail fast and route to fallback instead of continuing to hammer the degraded provider.
Single-layer retries prevent multiplicative explosion. If your application calls a service that calls another service, retries at every hop multiply. Three retries at each layer of a five-service chain produce 3^5 = 243 backend calls for a single user request. Pick one layer for retries, usually the outermost application layer or the routing gateway.
When using a routing layer like TheRouter, configure retries at the router level rather than in application code. The router sees all provider endpoints and can make smarter decisions about when to retry versus when to fail over to an alternative route. This avoids the double-retry problem where both your application and the gateway retry the same failed request.
Circuit breaker pattern for LLM providers
Retries handle transient failures. Circuit breakers handle systemic failures. The difference matters.
A circuit breaker monitors the failure rate over a rolling window and has three states:
Closed (normal operation): all requests pass through to the provider. The breaker tracks success/failure counts over the rolling window.
Open (tripped): the failure rate exceeded the threshold. All requests fail immediately without touching the provider. This gives the provider time to recover and prevents your application from wasting resources on requests that will fail.
Half-open (probing): after the cooldown period expires, the breaker allows a small number of test requests through. If they succeed, the breaker closes and normal traffic resumes. If they fail, the breaker reopens for another cooldown cycle.
For LLM APIs, configure these thresholds:
| Parameter | Recommended value | Rationale |
|---|---|---|
| Rolling window | 60 seconds | Long enough to detect sustained issues, short enough to react quickly |
| Failure threshold | 20-30% of requests | Higher than typical microservices because LLM APIs have baseline error rates |
| Consecutive failures to trip | 5-10 | Alternative trigger for low-traffic endpoints |
| Cooldown period | 30-60 seconds | Enough time for provider recovery |
| Half-open probe count | 1-3 requests | Minimal traffic to test recovery |
| Reset on success | After 2-3 consecutive half-open successes | Confirm recovery is stable, not a single lucky request |
LLM-specific circuit breaker triggers beyond standard HTTP error rates:
- Latency degradation: trip when p95 latency exceeds 3x the baseline. A provider returning 200 OK but taking 90 seconds per request is effectively degraded.
- Cost anomaly: trip when cost per request exceeds a configured threshold. This catches runaway agent loops and unexpected billing spikes.
- Streaming interruption rate: trip when more than 30% of streaming responses disconnect before the final chunk.
// Conceptual circuit breaker for LLM provider routing
interface CircuitBreakerConfig {
windowMs: number; // Rolling window: 60000
failureThreshold: number; // Percentage: 0.25
cooldownMs: number; // Cooldown: 30000
halfOpenProbes: number; // Probes: 2
latencyThresholdMs: number; // P95 cap: 45000
}
// When the breaker trips, route to the next provider in the
// fallback chain instead of returning an error to the user
function onCircuitOpen(provider: string, chain: string[]) {
const next = chain.find(p => getCircuitState(p) === 'closed');
if (next) {
routeTo(next); // Transparent fallback
} else {
serveDegraded(); // All providers degraded
}
}
Graceful degradation chains
When circuit breakers trip, your application needs somewhere to go. Graceful degradation chains define the fallback path from your primary model through progressively cheaper or simpler alternatives, down to cached responses or user-visible degraded mode.
A practical degradation chain for a coding assistant:
| Level | Model | Trigger | User impact |
|---|---|---|---|
| L0 (primary) | DeepSeek V4 Pro | Normal operation | Full capability |
| L1 (fallback) | DeepSeek V4 Flash | L0 circuit open or latency >45s | Slightly less accurate, much faster |
| L2 (budget) | Qwen3.8 Flash via DashScope | L0 + L1 circuits open | Different model, similar capability tier |
| L3 (cached) | Cached responses for common queries | All live providers degraded | Stale but available |
| L4 (degraded) | Error message with ETA | All fallbacks exhausted | Honest about unavailability |
The key design decisions:
Cross-provider fallback is more resilient than same-provider fallback. If OpenAI is down, falling back to another OpenAI model may not help. Falling back to Anthropic or DeepSeek via TheRouter's model fallback routing routes around provider-level outages.
Model downgrade is often better than provider outage. A user getting a response from a smaller model in 2 seconds is better served than waiting 30 seconds for a timeout from the large model. Configure downgrade thresholds based on latency, not just errors.
Cached responses need a freshness policy. For FAQ-style queries, a 1-hour cache is often acceptable. For real-time data questions, caching is not appropriate. Tag your queries with cacheability metadata.
Honest degradation builds trust. When all fallbacks fail, tell the user what happened and when you expect recovery. A clear "Our AI service is temporarily degraded; we expect recovery within 15 minutes" is better than a spinner that never resolves.
Rate limit handling and absorption
Rate limits are not errors; they are flow control. Production systems should absorb rate limits gracefully rather than treating them as failures.
Dual-axis tracking: LLM providers rate-limit on both RPM (requests per minute) and TPM (tokens per minute) simultaneously. You can stay within RPM while exceeding TPM, especially with long-context requests. Track both at your application layer, not just at the provider edge.
Pre-request token estimation prevents surprise TPM overruns. Use a tokenizer (tiktoken for OpenAI-compatible APIs, provider-specific libraries elsewhere) to estimate token count before sending the request. If the estimated tokens would exceed your remaining TPM budget, queue or shed the request rather than sending it and getting a 429.
Always set max_tokens to cap output length. Without this, a model generating an unusually long response can silently exhaust your TPM budget on a single request.
Burst absorption with queuing: instead of shedding requests that exceed your rate limit, queue them with a bounded wait time. A Redis-backed or in-memory priority queue with a 5-10 second maximum wait smooths burst traffic without losing requests:
import asyncio
from collections import deque
class TokenBucketLimiter:
def __init__(self, rpm: int, tpm: int):
self.rpm = rpm
self.tpm = tpm
self.request_tokens = rpm
self.token_tokens = tpm
self.queue: deque = deque()
async def acquire(self, estimated_tokens: int, timeout: float = 5.0):
"""Wait up to timeout seconds for capacity."""
deadline = asyncio.get_event_loop().time() + timeout
while asyncio.get_event_loop().time() < deadline:
if self.request_tokens > 0 and self.token_tokens >= estimated_tokens:
self.request_tokens -= 1
self.token_tokens -= estimated_tokens
return True
await asyncio.sleep(0.1)
return False # Shed or route to fallback
Streaming reliability
Streaming (SSE) responses add reliability challenges that batch requests do not have:
Partial response handling: a streaming response can disconnect midway. Your application should track how much output was received and whether it includes a completion marker (the [DONE] event in OpenAI-compatible APIs, or message_stop in Anthropic's format). A partial response without the stop marker needs to be flagged as incomplete.
Reconnection strategy: do not blindly reconnect and re-send the same prompt. The provider may have already processed part of the request and billed for it. Instead, mark the partial response as incomplete and either present it to the user with a warning or start a new request from scratch with adjusted context.
Per-chunk timeout: set a deadline for each SSE chunk, not just for the first byte. A provider that sends the first chunk in 2 seconds but then stalls for 60 seconds between chunks is effectively degraded. A 15-30 second per-chunk deadline catches this.
Error events mid-stream: both OpenAI and Anthropic can return error events after the initial 200 response. Your SSE parser must handle error event types that arrive mid-stream, not just HTTP-level error responses.
Health checks and proactive routing
Reactive resilience (retry after failure, trip after errors) is the baseline. Proactive resilience (route away from degraded providers before they affect users) is the next level.
Synthetic health probes: send lightweight test requests to each provider on a schedule (every 30-60 seconds). Use a small, fast model and a short prompt. Track response time and success rate. If a provider's probe latency exceeds 2x baseline or probes start failing, reduce its routing weight before user traffic is affected.
Provider status pages: monitor official status pages (status.openai.com, status.claude.com, status.deepseek.com) programmatically. When a provider reports a degraded or major outage, pre-emptively route traffic away.
Response header monitoring: OpenAI returns x-ratelimit-remaining-requests and x-ratelimit-remaining-tokens headers on every response. Use these to predict when you will hit limits and proactively throttle or reroute before getting a 429.
When using TheRouter as your routing layer, health-check routing is built into the model fallback configuration. The router monitors provider health across all traffic passing through it and routes around degraded endpoints without requiring application-level health check code.
Monitoring and SLO definition
You cannot harden what you do not measure. Define SLOs for your LLM API integration and alert on breaches.
Key metrics to track:
| Metric | Measurement | SLO example |
|---|---|---|
| Request success rate | 1 - (error_count / total_count) | >99.5% over 5-minute windows |
| Latency p50 | Median end-to-end request time | <5s for flash models, <15s for flagship |
| Latency p95 | 95th percentile request time | <15s for flash, <45s for flagship |
| Latency p99 | 99th percentile request time | <30s for flash, <90s for flagship |
| Circuit breaker open time | Total seconds breakers are open per hour | <300s/hour |
| Cost per request | Total spend / total requests | Below budget threshold |
| Retry rate | retry_count / total_count | <5% sustained, <10% peak |
| Streaming completion rate | complete_streams / started_streams | >99% |
Alert on:
- Success rate drops below SLO for 2 consecutive windows
- Any circuit breaker transitions to open
- Cost per request exceeds 2x the 7-day rolling average
- Retry rate exceeds 10% for more than 5 minutes
Putting it all together
The defense layers interact. Here is how they compose:
- Request arrives at your application
- Pre-request check: token estimation, rate limit budget check. If over budget, queue or shed.
- Route selection: health-check-driven routing picks the best available provider/model
- Timeout enforcement: connect + read + total timeouts wrap the API call
- Failure handling: if the request fails, classify the error
- Retry decision: retryable errors go through exponential backoff with jitter, within the retry budget
- Circuit breaker check: if the provider's circuit is open, skip to fallback immediately
- Fallback routing: if retries are exhausted or the circuit is open, route to the next model in the degradation chain
- Degraded mode: if all providers are down, serve cached responses or return an honest error
- Telemetry: log every decision point for monitoring and SLO tracking
This is not a recommendation to build all of this from scratch. An LLM routing layer handles steps 3, 4, 6, 7, and 8 at the infrastructure level, letting your application code focus on steps 1, 2, 5, 9, and 10.
FAQ
What timeout should I set for extended thinking / reasoning models?
Reasoning models like DeepSeek R1 or Claude with extended thinking can take 60-180 seconds to produce output. Set read timeouts to at least 120s for these models. Use streaming to get partial visibility into progress. If you use a total timeout, set it to 3x the expected generation time for the model and prompt length.
Should I retry tool calls and function calls?
Only if the tool execution sink is idempotent. If a tool call sends an email, creates a ticket, or writes to a database, retrying the request can duplicate the side effect. Check whether the previous call executed before retrying. Use idempotency keys when the provider supports them.
How do I handle rate limits across multiple API keys?
Track RPM and TPM per key, not per application. If you distribute traffic across multiple keys (for higher aggregate throughput), each key has its own limits. Your rate limiter needs to track each key independently and route to keys with remaining capacity.
What is the difference between application-level and gateway-level retries?
Application-level retries happen in your code. Gateway-level retries happen in the routing layer (like TheRouter or a reverse proxy). Use one or the other, not both. If both retry, you get multiplicative retry storms. Gateway-level retries are generally preferable because the gateway sees all traffic and can make smarter decisions about provider health.
How often should circuit breaker thresholds be tuned?
Review thresholds monthly. Provider reliability characteristics change over time. A threshold that was appropriate when a provider had 99.5% uptime may be too sensitive if the provider improves to 99.9%, or too loose if it degrades. Use your monitoring data to tune: if the breaker trips more than once a day on false positives, raise the threshold. If users report degraded experiences before the breaker trips, lower it.
When should I use a routing layer versus building retry and fallback logic in my application?
If you call one provider with one model, application-level retry logic is fine. If you call multiple providers, use multiple models, or need circuit breakers and health-check routing, a dedicated routing layer removes significant complexity from your application code and centralizes the resilience logic where it can be monitored and tuned independently.
Sources cited in this guide:
- OpenAI rate limits documentation (retrieved 2026-09-22)
- Anthropic Claude API errors (retrieved 2026-09-22)
- DeepSeek error codes (retrieved 2026-09-22)
- LLM API Resilience in Production — tianpan.co (retrieved 2026-09-22)
- Retries, fallbacks, and circuit breakers in LLM apps — Portkey (retrieved 2026-09-22)
- Circuit Breaker Pattern — AWS Prescriptive Guidance (retrieved 2026-09-22)