DeepSeek V4 Reasoning Effort Lands in Three API Formats: Why Your Proxy Is Probably Getting It Wrong

DeepSeek V4-Pro and V4-Flash now expose three-tier reasoning effort control across OpenAI, Anthropic, and Responses API formats. The catch: each format uses a different field, and proxies that strip extra_body silently override your effort to max.

Published via DeepSeek

Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

Diagram showing DeepSeek V4 reasoning effort API format differences between OpenAI, Anthropic, and Responses API

When DeepSeek launched V4-Pro GA on August 13, the headline story was peak/off-peak pricing. The more operationally significant change — quietly documented in the API docs — was the promotion of thinking effort control from an experimental option to a first-class, three-tier API parameter. And the implementation details matter for any team routing through a proxy or gateway.

What changed in the V4-Pro GA release

DeepSeek V4-Pro and V4-Flash now support three named effort levels for thinking mode: low, high, and max. The default, when nothing is specified, is high. This replaces the earlier binary on/off thinking toggle that many teams were using.

The important detail is how you actually set that effort — and it differs depending on which API format you're talking to.

OpenAI Chat Completions format:

response = client.chat.completions.create(
    model="deepseek-v4-pro",
    messages=[...],
    reasoning_effort="high",          # top-level field
    extra_body={"thinking": {"type": "enabled"}}  # required in extra_body
)

You need both: reasoning_effort at the top level to set effort, and {"thinking": {"type": "enabled"}} inside extra_body to actually activate thinking mode. If you omit extra_body, thinking is not enabled regardless of what reasoning_effort says.

Anthropic Messages format:

{
  "model": "deepseek-v4-pro",
  "thinking": {"type": "enabled"},
  "output_config": {"effort": "high"}
}

Responses API format:

{
  "model": "deepseek-v4-pro",
  "reasoning": {"effort": "high"}
}

Three formats, three different field paths to the same underlying capability.

Why proxies get this wrong silently

The problem is what happens when a proxy or gateway doesn't know about extra_body. Most OpenAI-compatible proxy implementations forward the standard Chat Completions fields — model, messages, temperature, max_tokens — and drop unknown top-level extensions. extra_body is a client-side SDK construct, not a wire-level field: the Python SDK merges it into the request body at call time. What actually goes over the wire is a regular JSON body with thinking at the top level alongside messages.

If your proxy normalizes or rebuilds the request body rather than forwarding it verbatim, it may strip thinking from the payload. The result: thinking mode is disabled, and your reasoning_effort setting becomes meaningless. You won't get an error. You'll get a normal response, billed at thinking-disabled rates, with no chain-of-thought output.

The Reddit community surfaced this exact pattern when testing through OpenCode: when the framework didn't pass the thinking block through, DeepSeek appeared to apply a server-side default — max effort in some versions, no thinking in others — regardless of the reasoning_effort value in the request.

The tool call context rule changes by effort mode

There's a second behavior that affects multi-turn agent pipelines: reasoning_content handling in tool call sequences.

When the model makes a tool call, the intermediate assistant's reasoning_content must be included when you concatenate context for the next turn. If the model did not make a tool call, the reasoning_content from that turn should be dropped before concatenating — passing it back is safe but adds unnecessary tokens.

This matters for agent frameworks that automatically strip assistant-side fields they don't recognize before building the next request. If your framework drops reasoning_content in tool call sequences, the model may behave inconsistently across turns. The effort level you chose — low, high, or max — doesn't change this rule; it applies the same way across all three.

What the Codex integration spec reveals

DeepSeek published a Codex integration config that exposes the internal routing behavior. The JSON spec for deepseek-v4-flash in Codex shows:

{
  "default_reasoning_level": "high",
  "supported_reasoning_levels": [
    {"effort": "low", "description": "Fast responses with lighter reasoning"},
    {"effort": "high", "description": "Extra high reasoning depth for complex problems"},
    {"effort": "max", "description": "Maximum reasoning depth for the hardest problems"}
  ]
}

This confirms that when you route V4-Flash through a Codex-aware client with no explicit effort setting, you get high by default. If your gateway is setting reasoning_effort=low in a header or config to save cost on simple tasks, and the Codex client is overriding it with its own default_reasoning_level after the fact, you have a conflict that won't surface as an error.

Cross-provider context: where DeepSeek's implementation differs

The three-tier effort system — low, high, max — mirrors what Anthropic offers for Claude's extended thinking budget. OpenAI's Fast mode and service tiers (standard, priority, fast) are orthogonal to reasoning depth and don't map directly. The significant difference with DeepSeek is that high is the default, not low, meaning the baseline cost for agent workloads running through V4-Pro is already at a mid-high reasoning budget.

For teams doing high-volume, low-complexity tasks (classification, extraction, simple generation) through V4-Flash, setting reasoning_effort=low can meaningfully reduce cost and latency — but only if your proxy actually passes thinking and reasoning_effort through correctly.

The routing decision for operators

Three questions to answer before you ship production traffic:

  1. Does your proxy forward the thinking block verbatim? Test by logging the actual request body your proxy sends to DeepSeek, not just what your SDK call looks like.

  2. Are you setting effort per-request or per-route? For mixed workloads — agents that need high or max, batch jobs that can use low — per-request effort control requires that your routing layer doesn't collapse effort to a single global default.

  3. Do your agent frameworks preserve reasoning_content in tool call sequences? Check your framework's message building logic. If it reconstructs assistant messages from a subset of response fields, it may silently drop reasoning_content on tool call turns.

What TheRouter users should watch

For teams routing to DeepSeek through TheRouter: the proxy layer passes request bodies verbatim to the upstream provider. The thinking and reasoning_effort fields will reach DeepSeek as long as your client sends them correctly. The risk is at the client side — verify your SDK version is sending the extra_body merge correctly if you're on the OpenAI format path.

If you're running mixed workloads and want to cap DeepSeek reasoning cost at the route level, routing simple tasks to deepseek-v4-flash with explicit reasoning_effort=low is the right policy — Flash's benchmarks on routine tasks are strong enough that low effort rarely loses meaningfully against high.

Help & contact