OpenAI-Compatible Coding Agent Setup: Cursor, Claude Code, Codex, and Zed with a Custom LLM Router
End-to-end guide for routing Cursor, Claude Code, OpenAI Codex CLI, Zed, and Devin Desktop through an OpenAI-compatible LLM router. Covers env vars, config files, model mapping, fallback chains, cost monitoring, and common gotchas.
Every coding agent ships with a default API endpoint. That works fine until your team needs cost caps, audit logs, provider fallback, or the ability to swap models without touching every developer's machine. An OpenAI-compatible LLM router sits between the coding agent and the model provider, handling all of that in one place.
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
This guide walks through the exact configuration for five major coding agents: Cursor, Claude Code, OpenAI Codex CLI, Zed, and Devin Desktop. Each section is self-contained — skip to the tool you use.
Quick-Reference Table
| Coding Agent | Config Method | Key Setting | Auth Variable |
|---|---|---|---|
| Cursor | GUI Settings | Override OpenAI Base URL | OpenAI API Key field |
| Claude Code | Env var / settings.json | ANTHROPIC_BASE_URL | ANTHROPIC_API_KEY or ANTHROPIC_AUTH_TOKEN |
| Codex CLI | config.toml | openai_base_url or model_provider block | OPENAI_API_KEY or provider-specific env var |
| Zed | Agent Settings JSON | language_models → OpenAI-compatible | Provider-ID-based _API_KEY env var |
| Devin Desktop | GUI Settings | Http: Proxy | Inherits from proxy |
Why Route Coding Agent Traffic Through a Gateway
Coding agents generate significant API spend — a single Codex session can burn through thousands of tokens per task. Routing through a gateway gives you:
- Cost visibility. See per-developer, per-model, per-project spend in one dashboard instead of checking five provider consoles.
- Provider fallback. When OpenAI returns 429 or 503, the router automatically tries your backup provider. We route OpenAI-compatible requests through configured providers and support provider/model routing and fallback when live product paths support it.
- Unified billing. One invoice instead of separate accounts with OpenAI, Anthropic, DeepSeek, and others. We provide unified billing/accounting surfaces where implemented.
- Governance. Rate limits, model allowlists, and spending caps enforced at the gateway level — developers never touch raw provider keys.
- Audit trail. Every request logged with user identity, model, tokens, and latency. Required for SOC 2 and enterprise compliance workflows.
Cursor
Cursor's settings GUI exposes two fields that redirect all OpenAI-format traffic through your router.
Step-by-Step
- Open Cursor Settings (
Cmd+,on macOS,Ctrl+,on Windows/Linux). - Navigate to Models → API Keys.
- Enter your router API key in the OpenAI API Key field.
- Check Override OpenAI Base URL.
- Enter your router endpoint URL (e.g.,
https://api.therouter.ai/v1). - Under Models, add the model IDs your router exposes (e.g.,
gpt-6-astra,claude-sonnet-4,deepseek-v4-pro).
Verification
Send a prompt in Cursor's chat or inline assistant. Check your router dashboard for the incoming request. If the request doesn't appear:
- Confirm the base URL ends with
/v1(Cursor appends/chat/completionsautomatically). - Make sure the API key is a valid key from your router, not from OpenAI directly.
- Check that the selected model ID matches what your router expects.
Gotchas
- Cursor's built-in models still use Cursor's own infrastructure. The override only applies when you select a model from your custom list. Cursor's default models (like
cursor-fast) bypass the base URL override entirely. - The API key field applies to all custom models. You cannot set per-model keys through the GUI. If you need per-model auth, handle it at the router level.
- Streaming is required. Cursor expects SSE streaming responses. Your router must support
stream: trueon/v1/chat/completions.
Claude Code
Claude Code uses environment variables to point at a custom gateway. The relevant variable is ANTHROPIC_BASE_URL, which tells Claude Code where to send its /v1/messages requests.
Step-by-Step
Option A: Shell environment variables (quick test)
export ANTHROPIC_BASE_URL=https://api.therouter.ai/api/anthropic
export ANTHROPIC_API_KEY=sk-your-router-key
claude
Option B: Settings file (persistent, recommended for teams)
Create or edit ~/.claude/settings.json:
{
"env": {
"ANTHROPIC_BASE_URL": "https://api.therouter.ai/api/anthropic",
"ANTHROPIC_API_KEY": "sk-your-router-key"
}
}
For organization-wide rollout, your admin can distribute the settings through a managed settings file that every developer's Claude Code picks up at startup.
Verification
- Run
claudeto start a session. - Type
/statusand check the Anthropic base URL line — it should show your router URL. - Send a test prompt. A normal response confirms routing works.
Gotchas
ANTHROPIC_BASE_URLmust point to the Anthropic-format path, not the OpenAI-compatible path. If your router exposes Anthropic-format at/api/anthropic, use that. The endpoint must serve/v1/messages.ANTHROPIC_AUTH_TOKENvsANTHROPIC_API_KEY: useANTHROPIC_AUTH_TOKENwhen your gateway expects a bearer token in theAuthorizationheader. UseANTHROPIC_API_KEYwhen it expectsx-api-key. When unsure, start withANTHROPIC_AUTH_TOKEN.- Background agents and supervisors may not inherit shell exports. Use the settings file approach for reliable routing in VS Code extensions and background sessions.
- Setting
ANTHROPIC_BASE_URLwithout a credential does not replace a claude.ai subscription — requests still route through your URL, but the subscription's billing and limits apply.
OpenAI Codex CLI
Codex CLI reads its configuration from ~/.codex/config.toml. The simplest approach is setting openai_base_url for the built-in OpenAI provider.
Step-by-Step
Option A: Built-in provider redirect
Edit ~/.codex/config.toml:
[model]
openai_base_url = "https://api.therouter.ai/v1"
Set your API key:
export OPENAI_API_KEY=sk-your-router-key
Option B: Custom model provider (for non-OpenAI models through the router)
[[model_provider]]
name = "therouter"
base_url = "https://api.therouter.ai/v1"
env_key = "THEROUTER_API_KEY"
wire_format = "openai"
[model_provider.models]
"deepseek-v4-pro" = { max_tokens = 131072 }
"claude-sonnet-4" = { max_tokens = 200000 }
Then set the environment variable:
export THEROUTER_API_KEY=sk-your-router-key
And select the model when launching Codex:
codex --model therouter/deepseek-v4-pro
Verification
Run codex and send a test task. Check your router logs for the incoming request. Codex uses the Responses API (/v1/responses) by default — verify your router supports it, or configure Codex to use the Chat Completions API if needed.
Gotchas
CODEX_HOMEoverrides the config directory. If set, Codex reads config from$CODEX_HOME/config.tomlinstead of~/.codex/config.toml.- Codex uses the Responses API by default, not Chat Completions. If your router only supports
/v1/chat/completions, check whether Codex offers awire_formator API mode toggle in your version. - Custom provider names must be unique. If you define a provider named
therouter, reference it astherouter/model-namein the--modelflag.
Zed
Zed supports OpenAI-compatible providers through its Agent Settings. Configuration lives in your settings.json file.
Step-by-Step
- Open Agent Settings: run agent: open settings from the command palette.
- Add an OpenAI-compatible provider under
language_models:
{
"language_models": {
"therouter": {
"type": "openai",
"api_url": "https://api.therouter.ai/v1",
"available_models": [
{
"name": "gpt-6-astra",
"display_name": "GPT-6 Astra (via TheRouter)",
"max_tokens": 200000
},
{
"name": "deepseek-v4-pro",
"display_name": "DeepSeek V4 Pro (via TheRouter)",
"max_tokens": 131072
},
{
"name": "claude-sonnet-4",
"display_name": "Claude Sonnet 4 (via TheRouter)",
"max_tokens": 200000
}
]
}
}
}
- Set the API key. Zed generates the env var name from the provider ID: provider
therouterreadsTHEROUTER_API_KEY.
export THEROUTER_API_KEY=sk-your-router-key
Alternatively, enter the key in Settings → AI → LLM Providers through the GUI, which stores it in your system keychain.
Verification
Open a new Zed Agent thread, select one of your custom models from the model picker, and send a test prompt. If the model doesn't appear, restart Zed after saving settings.json.
Gotchas
- Provider ID determines the env var name. The ID
my-gatewaybecomesMY_GATEWAY_API_KEY. Use simple, descriptive IDs. available_modelsis required for OpenAI-compatible providers. Zed cannot auto-discover models from arbitrary endpoints.- Zed Agent, Inline Assistant, and commit message generation all use the same LLM provider config. External Agents (like Claude Code running inside Zed's terminal) configure their own endpoints separately.
- Custom headers can be added via
custom_headersif your router requires extra auth headers beyond the API key.
Devin Desktop
Devin Desktop (formerly Windsurf) routes all LLM traffic through an HTTP proxy setting.
Step-by-Step
- Open Devin Desktop Settings (
Cmd+,on macOS,Ctrl+,on Windows/Linux). - Search for Http: Proxy.
- Enter your router URL:
https://api.therouter.ai. - Save and restart Devin Desktop.
Verification
Open the Devin Desktop chat panel and send a test message. Check your router dashboard for the incoming request.
Gotchas
- Devin Desktop uses a proxy model, not a base URL override. This means all HTTP traffic from the editor (not just LLM calls) may route through the proxy depending on your configuration.
- Authentication flows through the proxy. Your router must handle the auth headers that Devin Desktop sends to its upstream providers.
- Cascade is EOL. If you migrated from Windsurf, the legacy Cascade agent no longer functions. Use Devin's built-in agent instead.
Fallback Chain Setup
Once all agents route through your gateway, configure model fallback so coding workloads survive provider outages:
Primary: deepseek-v4-pro (lowest cost for coding tasks)
Fallback 1: claude-sonnet-4 (strong coding, higher cost)
Fallback 2: gpt-6-astra (broadest model, highest cost)
The router handles failover automatically — when the primary returns an error, the next model in the chain serves the request. Developers see a seamless experience regardless of which provider is healthy. We support provider/model routing and fallback when live product paths support it.
For model fallback configuration, set up the chain in your router dashboard or config file. Each coding agent sends requests to the same endpoint; the router decides which provider handles them.
Cost Monitoring
Routing through a single gateway gives you one place to track coding agent spend:
- Per-developer breakdown. Identify who generates the most tokens and which models they use.
- Per-model cost. Compare spend across
deepseek-v4-pro($0.70/M input) vsclaude-sonnet-4($3/M input) vsgpt-6-astrato find the best cost/quality balance. - Budget alerts. Set daily or monthly spending caps per team or per developer. The router can return 429 when a budget is exhausted instead of silently burning money.
We provide unified billing/accounting surfaces where implemented, so you get one invoice and one dashboard instead of reconciling five provider bills.
Common Gotchas and Troubleshooting
| Symptom | Likely Cause | Fix |
|---|---|---|
| "Model not found" error | Model ID in the agent doesn't match the router's model list | Check model IDs in your router config and the agent's settings |
| Requests bypass the router | Agent using built-in models instead of custom ones | Select a custom model from your model list |
| Auth errors (401/403) | Wrong API key or wrong auth header format | Verify the key belongs to your router, not the upstream provider |
| Streaming errors | Router doesn't support SSE streaming | Enable streaming support in your router; all coding agents require it |
| Slow first response | DNS resolution or TLS handshake to router | Ensure the router is in a region close to your developers |
| Codex "unsupported API" | Router lacks Responses API support | Check Codex docs for Chat Completions fallback option |
FAQ
Can I use different routers for different coding agents?
Yes. Each agent has its own configuration. You could route Cursor through one gateway and Claude Code through another. However, consolidating to a single router simplifies billing and monitoring.
Does routing through a gateway add latency?
Typically 5-20ms per request for a well-placed gateway. This is negligible for coding agent workloads where model inference itself takes 500ms-5s. The latency cost is outweighed by the reliability benefit of automatic failover.
Can I still use the agent's built-in models alongside routed models?
In most agents, yes. Cursor's built-in models, Claude Code's subscription access, and Zed's hosted models work independently of the custom endpoint configuration. You can switch between built-in and routed models per conversation.
What happens if my router goes down?
Without a secondary configuration, the coding agent fails to connect. To mitigate this, configure a secondary provider in agents that support it (like Codex's multiple model_provider blocks), or run a high-availability router setup.
Do I need separate router API keys per developer?
Recommended but not required. Per-developer keys enable usage attribution and individual revocation. Some routers support team-level keys with user identification through custom headers.
Does this work with self-hosted models?
Yes. If you run a model behind an OpenAI-compatible API (vLLM, Ollama, TGI), point your router at it as a provider, then point your coding agents at the router. The agents don't need to know the model is self-hosted.
Related Resources
- OpenAI-Compatible API Providers — full list of providers that work with this setup
- Cursor Custom API Endpoint Guide — deeper dive on Cursor-specific configuration
- AI Coding Agent API Routing Comparison — how routing compares across agents
- Coding Agent Model Routing Fall 2026 — latest model recommendations for coding workloads
- Model Fallback Configuration — detailed fallback chain setup
- OpenAI Provider — OpenAI-specific routing details
- Anthropic Provider — Anthropic-specific routing details