DeepSeek V4-Pro + Codex: Responses API Integration Guide for Coding Agents
How to wire DeepSeek V4-Pro and V4-Flash into Codex and other Responses-API-based coding agents. We walk through the setup, the compatibility surface, the gotchas, and the cost math — so you can decide whether DeepSeek belongs in your agentic coding stack.
DeepSeek's API now supports the Responses API format natively. That means Codex, OpenAI's agentic coding assistant, can talk to DeepSeek V4-Pro and V4-Flash without a compatibility shim. You point Codex at https://api.deepseek.com, register the models, and the agent starts working.
We tested this because coding-agent workloads hit the intersection of everything that matters in an API provider: tool calling reliability, streaming latency, thinking-mode depth, and cost per turn. DeepSeek V4-Pro scores well on all four, especially on cost. This guide covers the setup, the compatibility surface you can rely on, and the caveats you should know before pointing a production agent at it.
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
Getting Started in 3 Minutes
Step 1: Get your API key. Sign up at platform.deepseek.com and create a key from the API Keys page. Keys are shown once.
Step 2: Register DeepSeek in Codex. DeepSeek publishes an official Codex integration guide with a ready-made model configuration. The integration script registers deepseek-v4-flash and deepseek-v4-pro as available models in Codex, sets the base URL to https://api.deepseek.com, and configures Codex-specific parameters like apply_patch_tool_type, multi_agent_version, and truncation policy.
The short version: run the setup script DeepSeek provides, set your DEEPSEEK_API_KEY environment variable, and Codex will list both DeepSeek models alongside its built-in options.
Step 3: Send a smoke test. Before giving the agent write access to your repository, confirm that a basic Responses API call works:
from openai import OpenAI
client = OpenAI(
api_key="YOUR_DEEPSEEK_KEY",
base_url="https://api.deepseek.com",
)
response = client.responses.create(
model="deepseek-v4-flash",
instructions="You are a helpful coding assistant.",
input="Write a Python function that reverses a linked list.",
)
print(response.output_text)
If that returns a coherent answer, tool calls and streaming will work too. The Responses API on DeepSeek supports the same SSE event stream as OpenAI's implementation.
The Responses API on DeepSeek: What Works
DeepSeek's Responses API compatibility page is unusually transparent about what they support and what they don't. Here is the practical summary for coding-agent use cases:
Fully Supported
| Feature | Notes |
|---|---|
model | deepseek-v4-flash, deepseek-v4-pro, deepseek-v4-flash-vision-exp |
input | String or structured input item list |
instructions | Injected as first system message |
stream | Full SSE event stream with semantic events |
tools (function) | Standard function-calling tools |
tool_choice | none, auto, required, or a specific tool |
reasoning | effort parameter supported (low/high/max) |
temperature, top_p | Standard ranges (no effect in thinking mode) |
max_output_tokens | Up to 384K |
top_logprobs | Range 0–20 |
Partially Supported
| Feature | Status |
|---|---|
tools (web_search) | Supported, but code_interpreter and file_search are ignored |
text.format | Fully supported, but verbosity has no effect |
reasoning.summary | Accepted but no summary is generated |
parallel_tool_calls | Ignored; parallel tool calling is always enabled |
Not Supported
| Feature | Impact on Codex |
|---|---|
previous_response_id | DeepSeek's API is stateless; no server-side conversation history |
conversation | Same; multi-turn state must be managed client-side |
background | No background execution |
store | Responses always carry store: false |
truncation | Requests exceeding the context window return 400 |
For Codex specifically, the missing previous_response_id is a non-issue because Codex manages conversation state client-side. The lack of background mode means Codex runs all turns synchronously, which is the default behavior anyway.
Tool Calling for Coding Agents
Coding agents live and die by tool-call reliability. DeepSeek V4-Pro and V4-Flash support the full OpenAI function-calling schema through both the Chat Completions and Responses APIs. Both thinking and non-thinking modes support tool calls.
A typical coding-agent tool call flow looks like this:
tools = [
{
"type": "function",
"function": {
"name": "read_file",
"description": "Read the contents of a file at the given path.",
"parameters": {
"type": "object",
"properties": {
"path": {
"type": "string",
"description": "Absolute path to the file",
}
},
"required": ["path"],
},
},
},
]
response = client.responses.create(
model="deepseek-v4-pro",
instructions="You are a coding agent with file system access.",
input="Read the contents of /src/main.py and suggest improvements.",
tools=tools,
)
DeepSeek also supports a strict mode (via the /beta base URL for Chat Completions) that guarantees the model's tool-call output matches your JSON Schema exactly. For agent workflows where malformed tool calls cause cascading failures, this is worth testing, though the beta endpoint is separate from the Responses API path.
Codex-Specific Tool Types
Codex uses a proprietary apply_patch tool for writing file changes. DeepSeek's Codex integration configures this as "apply_patch_tool_type": "freeform", which means the model generates patch content in a freeform text format rather than structured JSON. The web_search_tool_type is set to "text", which maps to DeepSeek's native web search tool support.
Thinking Mode and Reasoning Effort
Both V4-Pro and V4-Flash default to thinking mode. For coding tasks that need deep reasoning (refactoring a complex codebase, debugging a subtle race condition), thinking mode produces noticeably better results. For straightforward tasks (renaming a variable, adding a docstring), it adds latency without proportional value.
The Responses API on DeepSeek accepts the reasoning.effort parameter with three levels:
| Effort | Behavior | Best For |
|---|---|---|
low | Fast, lighter reasoning | Simple edits, formatting, boilerplate |
high | Default; deeper chain-of-thought | Most coding tasks |
max | Maximum reasoning depth | Architecture decisions, complex debugging |
Codex maps these to its own reasoning level selector. The DeepSeek integration sets the default to high with all three levels available.
One important caveat: reasoning.summary is accepted but DeepSeek does not generate a reasoning summary. If your agent pipeline reads the summary to decide the next step, you won't get one from DeepSeek.
Pricing: The Cost Math for Coding Agents
This is where DeepSeek stands out. Coding agents are expensive because they generate long chains of tool calls, each with substantial input context. Here's how DeepSeek V4-Pro compares to the models Codex typically runs:
| Model | Input (Cache Miss) | Input (Cache Hit) | Output | Context |
|---|---|---|---|---|
| DeepSeek V4-Pro (off-peak) | $0.66 / 1M | $0.022 / 1M | $1.98 / 1M | 1M |
| DeepSeek V4-Pro (peak) | $1.32 / 1M | $0.044 / 1M | $3.96 / 1M | 1M |
| DeepSeek V4-Flash (off-peak) | $0.22 / 1M | $0.007 / 1M | $0.66 / 1M | 1M |
| DeepSeek V4-Flash (peak) | $0.44 / 1M | $0.014 / 1M | $1.32 / 1M | 1M |
Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday. All other hours are off-peak at 50% of peak rates.
The cache-hit pricing is the key number for coding agents. Agentic workflows repeatedly send the same system prompt, file contents, and conversation history. DeepSeek's automatic context caching means that after the first turn, much of your input hits the cache. At $0.022 per million tokens (off-peak, V4-Pro cache hit), you can run hundreds of agent turns for what a single uncached turn costs on some competing providers.
For cost-sensitive workloads, V4-Flash at off-peak cache-hit pricing ($0.007 / 1M input) is remarkably cheap. If your coding agent runs primarily on tasks that don't need Pro-level reasoning, Flash is the obvious default.
Rate Limits and Concurrency
DeepSeek imposes concurrency limits rather than RPM/TPM limits:
| Model | Concurrency Limit |
|---|---|
deepseek-v4-pro | 500 |
deepseek-v4-flash | 2,500 |
deepseek-v4-flash-vision-exp | 2,500 |
A request counts as one concurrent connection from send to response completion. Exceeding the limit returns HTTP 429.
For coding agents running in a team environment, the 500 concurrent connection limit on V4-Pro may be a constraint. Each developer running Codex with V4-Pro consumes one connection per active agent turn. If you have 50 developers each running multi-turn coding sessions simultaneously, you could approach the limit. For higher concurrency, DeepSeek offers capacity expansion at no additional cost through a request form.
Routing and Fallback Patterns
Running a coding agent on a single provider is risky. DeepSeek's API has experienced capacity-related slowdowns during peak periods. A routing layer that can fail over to an alternative provider protects your developers' workflow.
The pattern is straightforward because DeepSeek speaks the same OpenAI-compatible protocol. A model fallback configuration looks like:
# Example routing configuration
primary:
provider: deepseek
model: deepseek-v4-pro
fallback:
- provider: anthropic
model: claude-sonnet-5
- provider: openai
model: gpt-5.5-pro
When DeepSeek returns a 429 (concurrency exceeded) or 5xx (server error), the routing layer retries on the next provider. The coding agent never sees the failure because the request/response format is identical.
This also lets you do cost-aware routing. Route straightforward coding tasks to V4-Flash (cheapest), complex reasoning to V4-Pro, and fall back to Claude or GPT when DeepSeek is unavailable. The DeepSeek provider page on TheRouter has the current model lineup and routing status.
Smoke-Test Checklist Before Production
Before giving a coding agent backed by DeepSeek write access to your production codebase, verify each of these:
- Basic completion. Send a simple prompt and confirm you get a coherent response.
- Tool calls. Define a dummy tool, send a prompt that should trigger it, and verify the tool call is well-formed JSON.
- Streaming. Enable
stream=Trueand confirm you receiveresponse.output_text.deltaevents without gaps. - Thinking mode. Send a complex reasoning prompt with
reasoning={"effort": "max"}and verify the model produces chain-of-thought. - Long context. Send a request with a large file (50K+ tokens of code context) and confirm it doesn't truncate or error.
- Error handling. Intentionally exceed the concurrency limit and confirm your fallback logic catches the 429.
# Quick tool-call smoke test
response = client.responses.create(
model="deepseek-v4-pro",
instructions="You are a coding agent.",
input="What files are in the current directory?",
tools=[{
"type": "function",
"function": {
"name": "list_directory",
"description": "List files in a directory.",
"parameters": {
"type": "object",
"properties": {
"path": {"type": "string"}
},
"required": ["path"],
},
},
}],
)
# Verify the model emitted a tool call
for item in response.output:
if item.type == "function_call":
print(f"Tool call: {item.name}({item.arguments})")
Common Errors and Fixes
| Error | Cause | Fix |
|---|---|---|
400: input_image must have image_url or file_id | Image input without the required field | Use deepseek-v4-flash-vision-exp for image inputs; provide either image_url or file_id |
| 400: context length exceeded | Request exceeds the 1M token window | Truncate your conversation history; DeepSeek does not support the truncation parameter |
| 429: rate limit exceeded | Concurrency limit hit | Implement retry with exponential backoff; configure a fallback provider |
reasoning.summary returns empty | DeepSeek accepts but doesn't generate summaries | Don't depend on reasoning summaries for control flow; read the output directly |
previous_response_id not working | DeepSeek's API is stateless | Manage conversation state client-side (Codex already does this) |
When DeepSeek V4-Pro Belongs in Your Coding Stack
Pick DeepSeek V4-Pro as your coding-agent model when:
- Cost is a primary constraint. V4-Pro's off-peak pricing undercuts most frontier models significantly, and cache-hit rates on agentic workloads amplify the savings.
- You need deep reasoning but not the absolute frontier. V4-Pro's thinking mode handles complex refactors, debugging, and architecture questions well.
- You run during off-peak hours. If your team is in Asia-Pacific and works standard business hours, you naturally hit DeepSeek's off-peak pricing window (off-peak is everything except 01:00–04:00 and 06:00–10:00 UTC).
Pick V4-Flash when:
- Throughput matters more than depth. Flash's 2,500 concurrency limit and lower cost make it better for high-volume, less complex tasks.
- You're building a multi-model pipeline. Route simple tasks to Flash and complex tasks to Pro.
Consider a different provider when:
- You need background execution. DeepSeek doesn't support the
backgroundparameter. - You need server-side conversation state. The
previous_response_idandconversationparameters are not supported. - Your concurrency needs exceed 500 on Pro. Request a capacity increase, or use a routing layer to spread load across providers.
Further Reading
- DeepSeek API: The Complete Guide — Full API reference covering both Chat Completions and Responses API
- DeepSeek V4-Pro vs Flash: API Comparison — Detailed pricing and performance comparison
- Coding Agent Model Routing Comparison 2026 — How V4-Pro stacks up against Claude, Kimi K3, and GLM-5
- DeepSeek Provider Page — Current model lineup, routing status, and integration details
- DeepSeek V4-Pro Model Page — Specs, pricing, and benchmark data
- DeepSeek V4-Flash Model Page — Specs, pricing, and benchmark data