DeepSeek V4 Pro
How TheRouter serves this differently from the vendor
DeepSeek serves V4-Pro first-party at api.deepseek.com and api.deepseek.com/anthropic, with OpenAI Chat Completions and Anthropic-compatible request formats, 1M context, 384K maximum output, thinking enabled by default at high effort, and no Responses API support for Pro on the cited pricing page.
TheRouter exposes the same model through api.therouter.ai/v1/chat/completions as an OpenAI-compatible routed endpoint. Requests inherit TheRouter account routing, upstream selection, regional policy, tool-call normalisation, and TheRouter pricing rather than DeepSeek's first-party concurrency limits or first-party billing rates exactly.
API guide
Chat Completions
Standard OpenAI-compatible chat endpoint. Supports both non-thinking and thinking modes.
curl https://api.therouter.ai/v1/chat/completions \
-H "Authorization: Bearer $THEROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4-pro",
"messages": [{"role":"user","content":"Explain quantum error correction in 3 sentences."}]
}'Thinking / Reasoning mode
Enable chain-of-thought via extra_body. Returns reasoning_content separately from final answer.
curl https://api.therouter.ai/v1/chat/completions \
-H "Authorization: Bearer $THEROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4-pro",
"messages": [{"role":"user","content":"Solve: 9.11 vs 9.8 which is larger?"}],
"extra_body": {"thinking": {"type": "enabled"}},
"reasoning_effort": "max"
}'Tool calling
Native tool calling with full thinking-mode support. reasoning_content must be echoed back on subsequent turns when a tool call occurred.
curl https://api.therouter.ai/v1/chat/completions \
-H "Authorization: Bearer $THEROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4-pro",
"messages": [{"role":"user","content":"What is the weather in SF?"}],
"tools": [{"type":"function","function":{"name":"get_weather","parameters":{...}}}],
"tool_choice": "auto",
"extra_body": {"thinking": {"type": "enabled"}}
}'Part of
Recent coverage
- GuideLLM API Peak/Off-Peak Pricing: How to Schedule Workloads for Maximum Cost Savings
- GuideDeepSeek V4-Pro vs V4-Flash (August 2026): API Pricing, Benchmarks, and Routing Strategies
- GuideDeepSeek V4-Pro GA: Peak/Off-Peak Pricing, Responses API, and Routing Cost Optimization
- GuideQwen3.8-Max vs DeepSeek V4 Pro: Chinese Flagship Model API Comparison for Routing Operators
- GuideDeepSeek API Price Increase: What Routing Operators Need to Know and How to Prepare
- GuideStructured Output Across LLM API Providers: JSON Mode, JSON Schema, and What Each Provider Actually Supports (2026)
Fact ledger β every claim on this page traces here
| source | URL | retrieved | |
|---|---|---|---|
| Release date (preview) | api-docs.deepseek.com β | 2026-05-25 | verified |
| Architecture | api-docs.deepseek.com β | 2026-05-25 | verified |
| Context window | api-docs.deepseek.com β | 2026-05-25 | verified |
| Max output | api-docs.deepseek.com β | 2026-05-25 | verified |
| Reasoning modes | api-docs.deepseek.com β | 2026-05-25 | verified |
| License β code | huggingface.co β | 2026-05-25 | verified |
| License β weights | huggingface.co β | 2026-05-25 | verified |
| Open weights | huggingface.co β | 2026-05-25 | verified |
| Training tokens | api-docs.deepseek.com β | 2026-08-07 | unknown |
| Artificial Analysis Intelligence Index | api-docs.deepseek.com β | 2026-08-07 | unknown |
| LiveCodeBench Pass@1 | api-docs.deepseek.com β | 2026-08-07 | unknown |
| MMLU-Pro | api-docs.deepseek.com β | 2026-08-07 | unknown |
| GPQA Diamond | api-docs.deepseek.com β | 2026-08-07 | unknown |
| SWE-bench Verified | api-docs.deepseek.com β | 2026-08-07 | unknown |
| Terminal-Bench 2.0 | api-docs.deepseek.com β | 2026-08-07 | unknown |
| SimpleQA-Verified | api-docs.deepseek.com β | 2026-08-07 | unknown |
| BrowseComp | api-docs.deepseek.com β | 2026-08-07 | unknown |
| DeepSeek makes V4-Pro 75% discount permanent | api-docs.deepseek.com β | 2026-05-25 | verified |
| DeepSeek V4 series launches with 1M context and dual-tier MoE design | api-docs.deepseek.com β | 2026-05-25 | verified |
| When should I choose V4-Pro over V4-Flash? | api-docs.deepseek.com β | 2026-08-07 | to verify |
| Does V4-Pro support the same thinking mode API as V4-Flash? | api-docs.deepseek.com β | 2026-05-25 | to verify |