DeepSeek V4-Pro GA: Peak/Off-Peak Pricing, Responses API, and Routing Cost Optimization
DeepSeek V4-Pro went GA on August 13 with flexible reasoning effort, native Responses API support, and a new peak/off-peak pricing model effective August 16. We break down exactly what changed, how much it costs, and how to schedule workloads through your routing layer for maximum savings.
DeepSeek V4-Pro is now generally available. The GA release (model version DeepSeek-V4-Pro-0813) shipped on August 13, 2026 with three changes that matter if you route API traffic through DeepSeek: flexible reasoning effort across low/high/max levels, native OpenAI Responses API support optimized for Codex, and a new peak/off-peak pricing model that took effect at 16:00 UTC on August 16, 2026. Off-peak rates are exactly half of peak rates, which means your overnight batch jobs just got 50% cheaper — if you schedule them right.
This guide covers the confirmed pricing for both V4-Pro and V4-Flash, the peak/off-peak windows, how reasoning effort affects your cost profile, and how to use your routing layer to take advantage of time-of-day pricing.
What Changed on August 13
DeepSeek's V4-Pro GA announcement bundles several updates into one release:
-
V4-Pro goes GA. The model version is now
DeepSeek-V4-Pro-0813. The API model name remainsdeepseek-v4-pro— no code changes needed if you were already using the preview. -
Flexible reasoning effort. Both V4-Pro and V4-Flash now support
low,high, andmaxreasoning effort. The default ishigh. Uselowfor simple classification and extraction tasks where thinking overhead is wasted. Reservemaxfor complex multi-step problems where accuracy matters more than latency. -
Native Responses API. DeepSeek now supports the OpenAI Responses API format at the same
https://api.deepseek.combase URL. This is the first non-OpenAI provider to ship native Responses API compatibility, which means Codex and other Responses API consumers can connect directly with minimal configuration. -
Peak/off-peak pricing. Effective August 16, 2026 at 16:00 UTC, all DeepSeek API calls are billed at either peak or off-peak rates. Off-peak is 50% of peak.
Pricing: The Complete Table
All prices are per 1 million tokens. Data retrieved from the official DeepSeek pricing page on August 17, 2026.
V4-Flash (deepseek-v4-flash)
| Metric | Off-Peak | Peak |
|---|---|---|
| Input (cache hit) | $0.007 | $0.014 |
| Input (cache miss) | $0.22 | $0.44 |
| Output | $0.66 | $1.32 |
| Concurrency limit | 2,500 | 2,500 |
V4-Pro (deepseek-v4-pro)
| Metric | Off-Peak | Peak |
|---|---|---|
| Input (cache hit) | $0.022 | $0.044 |
| Input (cache miss) | $0.66 | $1.32 |
| Output | $1.98 | $3.96 |
| Concurrency limit | 500 | 500 |
Peak and Off-Peak Windows
Peak hours are 01:00–04:00 UTC and 06:00–10:00 UTC. All other hours are off-peak.
Translated to common time zones:
| Time Zone | Peak Windows | Off-Peak |
|---|---|---|
| UTC | 01:00–04:00, 06:00–10:00 | All other hours |
| Beijing (CST, UTC+8) | 09:00–12:00, 14:00–18:00 | All other hours |
| US Pacific (PDT, UTC-7) | 18:00–21:00, 23:00–03:00 | All other hours |
| US Eastern (EDT, UTC-4) | 21:00–00:00, 02:00–06:00 | All other hours |
| Central Europe (CEST, UTC+2) | 03:00–06:00, 08:00–12:00 | All other hours |
The peak windows correspond roughly to business hours in China (where the bulk of DeepSeek API traffic originates) plus early morning in Europe. If your user base is primarily in the Americas, most of your traffic will land in off-peak windows naturally.
Cost Comparison: Before and After
Before August 16, DeepSeek used flat pricing. Here is how the new rates compare for V4-Pro, which saw the most significant change:
| Scenario (V4-Pro, 1M tokens) | Old Flat Rate | New Off-Peak | New Peak |
|---|---|---|---|
| Input (cache miss) | $0.44 | $0.66 | $1.32 |
| Output | $1.32 | $1.98 | $3.96 |
The off-peak rate is 50% higher than the old flat rate, and the peak rate is 3x the old flat rate. This is a net price increase for unoptimized workloads — but operators who schedule batch processing during off-peak windows can still achieve costs comparable to the previous flat-rate structure.
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
Reasoning Effort: How to Pick the Right Level
V4-Pro and V4-Flash both default to thinking mode enabled at high effort. The effort levels map as follows:
| Requested Effort | Actual Mapped Effort |
|---|---|
low | Low — minimal chain-of-thought |
medium | High (maps to high internally) |
high | High (default) |
max | Max — deepest reasoning |
Setting reasoning_effort in the OpenAI SDK:
from openai import OpenAI
client = OpenAI(
api_key="YOUR_DEEPSEEK_KEY",
base_url="https://api.deepseek.com",
)
response = client.chat.completions.create(
model="deepseek-v4-pro",
messages=[{"role": "user", "content": "Analyze this contract clause..."}],
reasoning_effort="low",
extra_body={"thinking": {"type": "enabled"}},
)
To disable thinking entirely, set thinking.type to "disabled". This skips chain-of-thought completely and returns only the final answer — useful for simple tasks where you do not want to pay for reasoning tokens.
Cost implication: Thinking tokens are billed at the same rate as output tokens. Choosing low effort reduces the number of reasoning tokens the model generates, which directly reduces your cost per request. For classification or extraction tasks, low can cut thinking token usage by 60–80% compared to max.
Responses API: What It Means for Codex and Agent Workflows
DeepSeek's Responses API is a drop-in implementation of the OpenAI Responses API format. Supported features:
- Streaming with semantic SSE events —
response.output_text.delta,response.reasoning_text.delta, and typed tool call events - Function calling and web search via
toolsparameter - Reasoning effort via the
reasoning.effortfield instructionsparameter — inserted as the first system message
from openai import OpenAI
client = OpenAI(
api_key="YOUR_DEEPSEEK_KEY",
base_url="https://api.deepseek.com",
)
response = client.responses.create(
model="deepseek-v4-pro",
instructions="You are a helpful coding assistant.",
input="Refactor this function to use async/await.",
)
print(response.output_text)
Notable limitations compared to the OpenAI Responses API: previous_response_id, conversation, store, background, and metadata are not supported. DeepSeek's implementation is stateless — each request is independent. Unsupported parameters are silently ignored, so existing clients do not break.
Codex Integration
DeepSeek provides a one-click Codex setup with a model configuration that includes multi-agent v2 support, apply_patch tool type set to freeform, and a 1M token context window. The configuration supports both V4-Flash and V4-Pro models.
Routing DeepSeek V4-Pro Through TheRouter
TheRouter routes OpenAI-compatible requests through configured providers, which means you can use DeepSeek V4-Pro as a primary model and configure fallback routing to other providers when DeepSeek is unavailable.
A basic configuration that routes deepseek-v4-pro requests to DeepSeek directly:
from openai import OpenAI
client = OpenAI(
api_key="YOUR_THEROUTER_KEY",
base_url="https://therouter.ai/api/v1",
)
response = client.chat.completions.create(
model="deepseek/deepseek-v4-pro",
messages=[{"role": "user", "content": "Explain peak pricing"}],
)
Using DashScope as a Fallback Path
V4-Pro is also available on DashScope (Alibaba Cloud Bailian) under the model ID deepseek-v4-pro-0813. DashScope uses a different pricing structure (RMB-denominated, no peak/off-peak split as of this writing), which can serve as a cost-stabilizing fallback when DeepSeek's peak rates apply.
Scheduling Workloads for Off-Peak Savings
The peak/off-peak model creates a straightforward optimization strategy: defer deferrable work to off-peak windows.
Workloads that can be deferred:
- Batch evaluation runs
- Dataset labeling and annotation
- Document summarization pipelines
- Code review passes on non-urgent PRs
- Embedding generation for RAG indexes
Workloads that cannot be deferred:
- Interactive chat sessions
- Real-time coding assistant responses
- Latency-sensitive API endpoints
Implementation Approaches
1. Cron-scheduled batch jobs. Run your batch processing during off-peak hours. For a Beijing-based team, that means scheduling jobs outside 09:00–12:00 and 14:00–18:00 CST.
2. Queue-based deferral. Push non-urgent requests to a queue (SQS, RabbitMQ, Redis) and drain the queue during off-peak windows. This works well for document processing pipelines.
3. Provider-level Batch API. DeepSeek does not currently offer a Batch API equivalent to OpenAI's, but the off-peak pricing effectively serves the same purpose — you get a 50% discount for processing during low-demand windows, compared to OpenAI's 50% discount on Batch API calls.
Common Errors and Fixes
| Error | Cause | Fix |
|---|---|---|
| 400: reasoning_content must participate in context | Tool call in thinking mode — you did not pass back reasoning_content | Include reasoning_content from the assistant message in subsequent turns when tool calls are involved |
| 400: request exceeds context window | Input exceeds 1M tokens | Truncate input or split into multiple requests |
| Rate limit exceeded | Exceeded 500 concurrent requests (V4-Pro) or 2,500 (V4-Flash) | Add backoff logic, or route overflow to a secondary provider |
Production Checklist
Before going live with V4-Pro in a production environment:
- Confirm you are using the
deepseek-v4-promodel name (not legacydeepseek-chatordeepseek-reasoner) - Set
reasoning_effortexplicitly per use case — do not accept thehighdefault for simple tasks - Monitor peak vs. off-peak spend — if more than 60% of your token volume lands in peak windows, evaluate whether you can shift batch workloads
- Configure fallback routing to at least one alternative provider (V4-Flash, DashScope-hosted V4-Pro, or a different model family entirely)
- Test Responses API compatibility if you plan to use Codex or other Responses API consumers
- Set up alerting on the 500-request concurrency limit for V4-Pro
Further Reading
- DeepSeek V4-Pro GA Announcement
- DeepSeek Models and Pricing
- DeepSeek Thinking Mode
- DeepSeek Responses API
- DeepSeek Codex Integration
- DeepSeek API: The Complete Guide — our comprehensive guide covering the full V4 API surface
- DeepSeek Provider Page — model availability and routing configuration
- TheRouter Model Fallbacks — how to configure multi-provider fallback routing