← All articles

What Is an OpenAI-Compatible Router? The Developer's Guide to AI Gateways

An OpenAI-compatible router lets you point any OpenAI SDK at a single endpoint that routes requests across providers with fallback, cost tracking, and observability. This guide explains how it works, what to evaluate, and how to set one up in under a minute.

· TheRouter

Most teams start with a single LLM provider. The OpenAI SDK, one API key, a POST /v1/chat/completions call, and things work. Then requirements grow: a second provider for cost or latency reasons, a third as a safety net when the primary is rate-limited, and suddenly the codebase has provider-specific clients, separate error-handling paths, and three billing dashboards.

An OpenAI-compatible router solves this by sitting between your application and every provider. You keep using the OpenAI SDK. You change two lines of code: base_url and the API key. The router translates your request to whichever provider you choose, adds fallback when a model is unavailable, and gives you one bill and one set of logs.

This guide explains what that compatibility layer actually does, what to evaluate before choosing a router, and how to send your first routed request in under a minute.

OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.

How OpenAI compatibility works

The OpenAI Chat Completions API has become the de facto standard for LLM requests. The contract is straightforward: send a JSON body with model, messages, and optional parameters like temperature, max_tokens, tools, and stream to a /v1/chat/completions endpoint. Get back a JSON response with choices, usage, and metadata.

An OpenAI-compatible router accepts that exact contract. Under the hood it does three things:

  1. Model resolution. The model field carries a provider prefix, for example anthropic/claude-sonnet-4.5 or deepseek/deepseek-chat. The router maps this to the upstream provider's endpoint and authentication.

  2. Request translation. Most providers accept OpenAI-shaped requests, but edge cases differ: parameter names for reasoning effort, thinking-token budgets, multimodal input formats, tool-calling schemas. The router normalizes these before forwarding.

  3. Response normalization. The response comes back in the standard choices[0].message.content shape regardless of which provider served it. Provider-specific fields (like reasoning_content from DeepSeek) are preserved when verified.

The result: your application code never changes when you add a provider, swap a model, or set up fallback.

What a router adds beyond compatibility

Translating request formats is table stakes. The practical value of a routing layer shows up in five areas:

Fallback and retry

When a provider returns 429 (rate limit), 503 (overloaded), or times out, the router tries the next model in a fallback chain. Your application receives a successful response; the model field in the response tells you which provider actually served it.

Cost tracking

Every request flows through one endpoint. The router meters input tokens, output tokens, and cached tokens per model, per team, per API key. Instead of reconciling three vendor dashboards, you get one cost view.

Observability

Latency, error rates, token counts, and fallback events are logged in one place. This is the data you need for monitoring and alerting without instrumenting each provider client separately.

Access control

One API key per team or application, with per-key model allowlists, rate limits, and spend caps. The router handles upstream authentication with each provider's credentials, which stay on the server.

Streaming compatibility

SSE streaming works the same way regardless of provider. The router translates each provider's chunk format into the standard data: {"choices":[...]} shape. For details on streaming differences, see the cross-provider streaming guide.

Five things to check before choosing a router

Not all routers are equivalent. Here is what matters in production:

1. Compatibility depth. Does it handle tool calling, structured output (response_format), streaming, vision inputs, and reasoning-effort parameters across providers? Shallow compatibility breaks when you move past basic chat.

2. Latency overhead. Every proxy adds round-trip time. Measure the added latency at P50 and P99 for your traffic pattern. Good routers add single-digit milliseconds; bad ones add hundreds.

3. Pricing model. Some routers charge per request, some take a percentage markup on provider cost, some charge a flat platform fee. Calculate total cost at your expected volume, including the provider fees the router passes through. For a detailed comparison, see AI model router pricing.

4. Provider coverage. Check whether the router supports the specific models and providers you need today, and whether adding new ones requires configuration changes or code changes.

5. Observability surface. Logs, dashboards, and webhook alerts matter more than feature lists. If you cannot see which requests fell back, which models are slow, and where your spend is going, the router is not doing its job.

Quick-start: your first routed request in 60 seconds

This example uses TheRouter as the gateway. TheRouter routes OpenAI-compatible requests through configured providers and supports provider/model fallback.

Python

from openai import OpenAI

client = OpenAI(
    base_url="https://api.therouter.ai/v1",
    api_key="<THEROUTER_API_KEY>",
)

response = client.chat.completions.create(
    model="anthropic/claude-sonnet-4.5",
    messages=[{"role": "user", "content": "Explain what an LLM router does in two sentences."}],
)
print(response.choices[0].message.content)

Node.js

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.therouter.ai/v1",
  apiKey: "<THEROUTER_API_KEY>",
});

const completion = await client.chat.completions.create({
  model: "anthropic/claude-sonnet-4.5",
  messages: [{ role: "user", content: "Explain what an LLM router does in two sentences." }],
});
console.log(completion.choices[0].message.content);

cURL

curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-chat",
    "messages": [{"role": "user", "content": "What is an LLM router?"}]
  }'

Two things changed from a direct OpenAI call: the base URL and the API key. Everything else, including the SDK, the request shape, and the response format, stays the same.

Adding a fallback chain

The point of routing is reliability. Pass a models array to define a fallback chain. TheRouter attempts models in order and stops at the first successful response:

const completion = await client.chat.completions.create({
  model: "openai/gpt-5",
  extra_body: {
    models: ["anthropic/claude-sonnet-4.5", "deepseek/deepseek-chat"],
  },
  messages: [{ role: "user", content: "Classify this support ticket." }],
});

If gpt-5 returns a rate-limit error, the router retries with claude-sonnet-4.5. If that also fails, it falls back to deepseek-chat. The response model field tells you which one served the request. For the full fallback API, see the model fallbacks guide.

Common use cases

Multi-provider failover

The most common reason to use a router. Provider outages, rate limits, and moderation refusals are not hypothetical, they happen weekly. A fallback chain turns a user-facing error into a transparent retry.

Cost optimization

Route low-complexity requests to cheaper models and reserve expensive frontier models for tasks where quality justifies the cost. See the cost optimization routing guide for concrete bucketing strategies.

Coding agent routing

Tools like Cursor, Claude Code, and Windsurf accept a custom base_url. Point them at a router to control which models they use, enforce spend limits per developer, and log every request. The coding agent API routing comparison covers tool-by-tool configuration.

Model A/B testing

Split traffic between two models by request attribute and compare output quality, latency, and cost. The router logs both paths; your application code stays the same.

How TheRouter works as an OpenAI-compatible router

TheRouter routes OpenAI-compatible requests through configured providers and supports provider/model routing and fallback. It provides unified billing and accounting surfaces, and supports async media jobs through /v1/jobs/:id for media generation.

It is not the only option. OpenRouter, LiteLLM, Portkey, and Cloudflare AI Gateway all provide OpenAI-compatible routing with different trade-offs. Our gateway comparison covers the landscape.

What TheRouter does specifically:

  • Single endpoint: https://api.therouter.ai/v1 accepts any OpenAI SDK call
  • Provider-prefixed models: anthropic/claude-sonnet-4.5, deepseek/deepseek-chat, openai/gpt-5, dashscope/qwen3.8-max
  • Fallback chains: models array with priority-ordered alternatives
  • Streaming: SSE-compatible streaming across all routed providers
  • Tool calling: forwarded to providers that support it, with format translation where needed

FAQ

Does streaming work the same through a router? Yes. The router translates each provider's SSE chunk format into the standard OpenAI streaming shape. Your stream: true code works without changes.

What about tool calling and function calling? Tool calling is forwarded to the upstream provider. The router translates tool schemas when providers use different formats. See the function calling comparison for provider-specific differences.

Do provider-specific response fields survive? Fields like DeepSeek's reasoning_content are preserved when verified. The router does not strip provider extensions, but it does not guarantee every undocumented field either.

What happens if all models in my fallback chain fail? The router returns the error from the last attempted model. Your application receives a standard error response and can handle it the same way it would handle a direct provider error.

Is there added latency? Any proxy adds some latency. A well-implemented router adds single-digit milliseconds for the routing decision. The dominant factor is still the provider's inference time, which is typically hundreds of milliseconds to seconds.

Can I use this with coding agents like Cursor or Claude Code? Yes. Both accept a custom base URL. Point them at the router endpoint and they send requests through the routing layer. See the coding tools setup guide for step-by-step configuration.


Sources: OpenAI API reference, OpenRouter — LLM Gateway explainer, TheRouter quickstart, TheRouter model fallbacks, LiteLLM GitHub, Cloudflare AI Gateway docs

Help & contact