← All articles

OpenAI-Compatible Coding Agent Setup: Cursor, Claude Code, Codex, and Zed with a Custom LLM Router

End-to-end guide for routing Cursor, Claude Code, OpenAI Codex CLI, Zed, and Devin Desktop through an OpenAI-compatible LLM router. Covers env vars, config files, model mapping, fallback chains, cost monitoring, and common gotchas.

· TheRouter

Every coding agent ships with a default API endpoint. That works fine until your team needs cost caps, audit logs, provider fallback, or the ability to swap models without touching every developer's machine. An OpenAI-compatible LLM router sits between the coding agent and the model provider, handling all of that in one place.

OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.

This guide walks through the exact configuration for five major coding agents: Cursor, Claude Code, OpenAI Codex CLI, Zed, and Devin Desktop. Each section is self-contained — skip to the tool you use.

Quick-Reference Table

Coding AgentConfig MethodKey SettingAuth Variable
CursorGUI SettingsOverride OpenAI Base URLOpenAI API Key field
Claude CodeEnv var / settings.jsonANTHROPIC_BASE_URLANTHROPIC_API_KEY or ANTHROPIC_AUTH_TOKEN
Codex CLIconfig.tomlopenai_base_url or model_provider blockOPENAI_API_KEY or provider-specific env var
ZedAgent Settings JSONlanguage_models → OpenAI-compatibleProvider-ID-based _API_KEY env var
Devin DesktopGUI SettingsHttp: ProxyInherits from proxy

Why Route Coding Agent Traffic Through a Gateway

Coding agents generate significant API spend — a single Codex session can burn through thousands of tokens per task. Routing through a gateway gives you:

  • Cost visibility. See per-developer, per-model, per-project spend in one dashboard instead of checking five provider consoles.
  • Provider fallback. When OpenAI returns 429 or 503, the router automatically tries your backup provider. We route OpenAI-compatible requests through configured providers and support provider/model routing and fallback when live product paths support it.
  • Unified billing. One invoice instead of separate accounts with OpenAI, Anthropic, DeepSeek, and others. We provide unified billing/accounting surfaces where implemented.
  • Governance. Rate limits, model allowlists, and spending caps enforced at the gateway level — developers never touch raw provider keys.
  • Audit trail. Every request logged with user identity, model, tokens, and latency. Required for SOC 2 and enterprise compliance workflows.

Cursor

Cursor's settings GUI exposes two fields that redirect all OpenAI-format traffic through your router.

Step-by-Step

  1. Open Cursor Settings (Cmd+, on macOS, Ctrl+, on Windows/Linux).
  2. Navigate to Models → API Keys.
  3. Enter your router API key in the OpenAI API Key field.
  4. Check Override OpenAI Base URL.
  5. Enter your router endpoint URL (e.g., https://api.therouter.ai/v1).
  6. Under Models, add the model IDs your router exposes (e.g., gpt-6-astra, claude-sonnet-4, deepseek-v4-pro).

Verification

Send a prompt in Cursor's chat or inline assistant. Check your router dashboard for the incoming request. If the request doesn't appear:

  • Confirm the base URL ends with /v1 (Cursor appends /chat/completions automatically).
  • Make sure the API key is a valid key from your router, not from OpenAI directly.
  • Check that the selected model ID matches what your router expects.

Gotchas

  • Cursor's built-in models still use Cursor's own infrastructure. The override only applies when you select a model from your custom list. Cursor's default models (like cursor-fast) bypass the base URL override entirely.
  • The API key field applies to all custom models. You cannot set per-model keys through the GUI. If you need per-model auth, handle it at the router level.
  • Streaming is required. Cursor expects SSE streaming responses. Your router must support stream: true on /v1/chat/completions.

Claude Code

Claude Code uses environment variables to point at a custom gateway. The relevant variable is ANTHROPIC_BASE_URL, which tells Claude Code where to send its /v1/messages requests.

Step-by-Step

Option A: Shell environment variables (quick test)

export ANTHROPIC_BASE_URL=https://api.therouter.ai/api/anthropic
export ANTHROPIC_API_KEY=sk-your-router-key
claude

Option B: Settings file (persistent, recommended for teams)

Create or edit ~/.claude/settings.json:

{
  "env": {
    "ANTHROPIC_BASE_URL": "https://api.therouter.ai/api/anthropic",
    "ANTHROPIC_API_KEY": "sk-your-router-key"
  }
}

For organization-wide rollout, your admin can distribute the settings through a managed settings file that every developer's Claude Code picks up at startup.

Verification

  1. Run claude to start a session.
  2. Type /status and check the Anthropic base URL line — it should show your router URL.
  3. Send a test prompt. A normal response confirms routing works.

Gotchas

  • ANTHROPIC_BASE_URL must point to the Anthropic-format path, not the OpenAI-compatible path. If your router exposes Anthropic-format at /api/anthropic, use that. The endpoint must serve /v1/messages.
  • ANTHROPIC_AUTH_TOKEN vs ANTHROPIC_API_KEY: use ANTHROPIC_AUTH_TOKEN when your gateway expects a bearer token in the Authorization header. Use ANTHROPIC_API_KEY when it expects x-api-key. When unsure, start with ANTHROPIC_AUTH_TOKEN.
  • Background agents and supervisors may not inherit shell exports. Use the settings file approach for reliable routing in VS Code extensions and background sessions.
  • Setting ANTHROPIC_BASE_URL without a credential does not replace a claude.ai subscription — requests still route through your URL, but the subscription's billing and limits apply.

OpenAI Codex CLI

Codex CLI reads its configuration from ~/.codex/config.toml. The simplest approach is setting openai_base_url for the built-in OpenAI provider.

Step-by-Step

Option A: Built-in provider redirect

Edit ~/.codex/config.toml:

[model]
openai_base_url = "https://api.therouter.ai/v1"

Set your API key:

export OPENAI_API_KEY=sk-your-router-key

Option B: Custom model provider (for non-OpenAI models through the router)

[[model_provider]]
name = "therouter"
base_url = "https://api.therouter.ai/v1"
env_key = "THEROUTER_API_KEY"
wire_format = "openai"

[model_provider.models]
"deepseek-v4-pro" = { max_tokens = 131072 }
"claude-sonnet-4" = { max_tokens = 200000 }

Then set the environment variable:

export THEROUTER_API_KEY=sk-your-router-key

And select the model when launching Codex:

codex --model therouter/deepseek-v4-pro

Verification

Run codex and send a test task. Check your router logs for the incoming request. Codex uses the Responses API (/v1/responses) by default — verify your router supports it, or configure Codex to use the Chat Completions API if needed.

Gotchas

  • CODEX_HOME overrides the config directory. If set, Codex reads config from $CODEX_HOME/config.toml instead of ~/.codex/config.toml.
  • Codex uses the Responses API by default, not Chat Completions. If your router only supports /v1/chat/completions, check whether Codex offers a wire_format or API mode toggle in your version.
  • Custom provider names must be unique. If you define a provider named therouter, reference it as therouter/model-name in the --model flag.

Zed

Zed supports OpenAI-compatible providers through its Agent Settings. Configuration lives in your settings.json file.

Step-by-Step

  1. Open Agent Settings: run agent: open settings from the command palette.
  2. Add an OpenAI-compatible provider under language_models:
{
  "language_models": {
    "therouter": {
      "type": "openai",
      "api_url": "https://api.therouter.ai/v1",
      "available_models": [
        {
          "name": "gpt-6-astra",
          "display_name": "GPT-6 Astra (via TheRouter)",
          "max_tokens": 200000
        },
        {
          "name": "deepseek-v4-pro",
          "display_name": "DeepSeek V4 Pro (via TheRouter)",
          "max_tokens": 131072
        },
        {
          "name": "claude-sonnet-4",
          "display_name": "Claude Sonnet 4 (via TheRouter)",
          "max_tokens": 200000
        }
      ]
    }
  }
}
  1. Set the API key. Zed generates the env var name from the provider ID: provider therouter reads THEROUTER_API_KEY.
export THEROUTER_API_KEY=sk-your-router-key

Alternatively, enter the key in Settings → AI → LLM Providers through the GUI, which stores it in your system keychain.

Verification

Open a new Zed Agent thread, select one of your custom models from the model picker, and send a test prompt. If the model doesn't appear, restart Zed after saving settings.json.

Gotchas

  • Provider ID determines the env var name. The ID my-gateway becomes MY_GATEWAY_API_KEY. Use simple, descriptive IDs.
  • available_models is required for OpenAI-compatible providers. Zed cannot auto-discover models from arbitrary endpoints.
  • Zed Agent, Inline Assistant, and commit message generation all use the same LLM provider config. External Agents (like Claude Code running inside Zed's terminal) configure their own endpoints separately.
  • Custom headers can be added via custom_headers if your router requires extra auth headers beyond the API key.

Devin Desktop

Devin Desktop (formerly Windsurf) routes all LLM traffic through an HTTP proxy setting.

Step-by-Step

  1. Open Devin Desktop Settings (Cmd+, on macOS, Ctrl+, on Windows/Linux).
  2. Search for Http: Proxy.
  3. Enter your router URL: https://api.therouter.ai.
  4. Save and restart Devin Desktop.

Verification

Open the Devin Desktop chat panel and send a test message. Check your router dashboard for the incoming request.

Gotchas

  • Devin Desktop uses a proxy model, not a base URL override. This means all HTTP traffic from the editor (not just LLM calls) may route through the proxy depending on your configuration.
  • Authentication flows through the proxy. Your router must handle the auth headers that Devin Desktop sends to its upstream providers.
  • Cascade is EOL. If you migrated from Windsurf, the legacy Cascade agent no longer functions. Use Devin's built-in agent instead.

Fallback Chain Setup

Once all agents route through your gateway, configure model fallback so coding workloads survive provider outages:

Primary:    deepseek-v4-pro (lowest cost for coding tasks)
Fallback 1: claude-sonnet-4 (strong coding, higher cost)
Fallback 2: gpt-6-astra (broadest model, highest cost)

The router handles failover automatically — when the primary returns an error, the next model in the chain serves the request. Developers see a seamless experience regardless of which provider is healthy. We support provider/model routing and fallback when live product paths support it.

For model fallback configuration, set up the chain in your router dashboard or config file. Each coding agent sends requests to the same endpoint; the router decides which provider handles them.

Cost Monitoring

Routing through a single gateway gives you one place to track coding agent spend:

  • Per-developer breakdown. Identify who generates the most tokens and which models they use.
  • Per-model cost. Compare spend across deepseek-v4-pro ($0.70/M input) vs claude-sonnet-4 ($3/M input) vs gpt-6-astra to find the best cost/quality balance.
  • Budget alerts. Set daily or monthly spending caps per team or per developer. The router can return 429 when a budget is exhausted instead of silently burning money.

We provide unified billing/accounting surfaces where implemented, so you get one invoice and one dashboard instead of reconciling five provider bills.

Common Gotchas and Troubleshooting

SymptomLikely CauseFix
"Model not found" errorModel ID in the agent doesn't match the router's model listCheck model IDs in your router config and the agent's settings
Requests bypass the routerAgent using built-in models instead of custom onesSelect a custom model from your model list
Auth errors (401/403)Wrong API key or wrong auth header formatVerify the key belongs to your router, not the upstream provider
Streaming errorsRouter doesn't support SSE streamingEnable streaming support in your router; all coding agents require it
Slow first responseDNS resolution or TLS handshake to routerEnsure the router is in a region close to your developers
Codex "unsupported API"Router lacks Responses API supportCheck Codex docs for Chat Completions fallback option

FAQ

Can I use different routers for different coding agents?

Yes. Each agent has its own configuration. You could route Cursor through one gateway and Claude Code through another. However, consolidating to a single router simplifies billing and monitoring.

Does routing through a gateway add latency?

Typically 5-20ms per request for a well-placed gateway. This is negligible for coding agent workloads where model inference itself takes 500ms-5s. The latency cost is outweighed by the reliability benefit of automatic failover.

Can I still use the agent's built-in models alongside routed models?

In most agents, yes. Cursor's built-in models, Claude Code's subscription access, and Zed's hosted models work independently of the custom endpoint configuration. You can switch between built-in and routed models per conversation.

What happens if my router goes down?

Without a secondary configuration, the coding agent fails to connect. To mitigate this, configure a secondary provider in agents that support it (like Codex's multiple model_provider blocks), or run a high-availability router setup.

Do I need separate router API keys per developer?

Recommended but not required. Per-developer keys enable usage attribution and individual revocation. Some routers support team-level keys with user identification through custom headers.

Does this work with self-hosted models?

Yes. If you run a model behind an OpenAI-compatible API (vLLM, Ollama, TGI), point your router at it as a provider, then point your coding agents at the router. The agents don't need to know the model is self-hosted.

Related Resources

Help & contact