← All articles

DeepSeek V4-Pro GA: Peak/Off-Peak Pricing, Responses API, and Routing Cost Optimization

DeepSeek V4-Pro went GA on August 13 with flexible reasoning effort, native Responses API support, and a new peak/off-peak pricing model effective August 16. We break down exactly what changed, how much it costs, and how to schedule workloads through your routing layer for maximum savings.

· TheRouter

DeepSeek V4-Pro is now generally available. The GA release (model version DeepSeek-V4-Pro-0813) shipped on August 13, 2026 with three changes that matter if you route API traffic through DeepSeek: flexible reasoning effort across low/high/max levels, native OpenAI Responses API support optimized for Codex, and a new peak/off-peak pricing model that took effect at 16:00 UTC on August 16, 2026. Off-peak rates are exactly half of peak rates, which means your overnight batch jobs just got 50% cheaper — if you schedule them right.

This guide covers the confirmed pricing for both V4-Pro and V4-Flash, the peak/off-peak windows, how reasoning effort affects your cost profile, and how to use your routing layer to take advantage of time-of-day pricing.

What Changed on August 13

DeepSeek's V4-Pro GA announcement bundles several updates into one release:

  1. V4-Pro goes GA. The model version is now DeepSeek-V4-Pro-0813. The API model name remains deepseek-v4-pro — no code changes needed if you were already using the preview.

  2. Flexible reasoning effort. Both V4-Pro and V4-Flash now support low, high, and max reasoning effort. The default is high. Use low for simple classification and extraction tasks where thinking overhead is wasted. Reserve max for complex multi-step problems where accuracy matters more than latency.

  3. Native Responses API. DeepSeek now supports the OpenAI Responses API format at the same https://api.deepseek.com base URL. This is the first non-OpenAI provider to ship native Responses API compatibility, which means Codex and other Responses API consumers can connect directly with minimal configuration.

  4. Peak/off-peak pricing. Effective August 16, 2026 at 16:00 UTC, all DeepSeek API calls are billed at either peak or off-peak rates. Off-peak is 50% of peak.

Pricing: The Complete Table

All prices are per 1 million tokens. Data retrieved from the official DeepSeek pricing page on August 17, 2026.

V4-Flash (deepseek-v4-flash)

MetricOff-PeakPeak
Input (cache hit)$0.007$0.014
Input (cache miss)$0.22$0.44
Output$0.66$1.32
Concurrency limit2,5002,500

V4-Pro (deepseek-v4-pro)

MetricOff-PeakPeak
Input (cache hit)$0.022$0.044
Input (cache miss)$0.66$1.32
Output$1.98$3.96
Concurrency limit500500

Peak and Off-Peak Windows

Peak hours are 01:00–04:00 UTC and 06:00–10:00 UTC. All other hours are off-peak.

Translated to common time zones:

Time ZonePeak WindowsOff-Peak
UTC01:00–04:00, 06:00–10:00All other hours
Beijing (CST, UTC+8)09:00–12:00, 14:00–18:00All other hours
US Pacific (PDT, UTC-7)18:00–21:00, 23:00–03:00All other hours
US Eastern (EDT, UTC-4)21:00–00:00, 02:00–06:00All other hours
Central Europe (CEST, UTC+2)03:00–06:00, 08:00–12:00All other hours

The peak windows correspond roughly to business hours in China (where the bulk of DeepSeek API traffic originates) plus early morning in Europe. If your user base is primarily in the Americas, most of your traffic will land in off-peak windows naturally.

Cost Comparison: Before and After

Before August 16, DeepSeek used flat pricing. Here is how the new rates compare for V4-Pro, which saw the most significant change:

Scenario (V4-Pro, 1M tokens)Old Flat RateNew Off-PeakNew Peak
Input (cache miss)$0.44$0.66$1.32
Output$1.32$1.98$3.96

The off-peak rate is 50% higher than the old flat rate, and the peak rate is 3x the old flat rate. This is a net price increase for unoptimized workloads — but operators who schedule batch processing during off-peak windows can still achieve costs comparable to the previous flat-rate structure.

OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.

Reasoning Effort: How to Pick the Right Level

V4-Pro and V4-Flash both default to thinking mode enabled at high effort. The effort levels map as follows:

Requested EffortActual Mapped Effort
lowLow — minimal chain-of-thought
mediumHigh (maps to high internally)
highHigh (default)
maxMax — deepest reasoning

Setting reasoning_effort in the OpenAI SDK:

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_DEEPSEEK_KEY",
    base_url="https://api.deepseek.com",
)

response = client.chat.completions.create(
    model="deepseek-v4-pro",
    messages=[{"role": "user", "content": "Analyze this contract clause..."}],
    reasoning_effort="low",
    extra_body={"thinking": {"type": "enabled"}},
)

To disable thinking entirely, set thinking.type to "disabled". This skips chain-of-thought completely and returns only the final answer — useful for simple tasks where you do not want to pay for reasoning tokens.

Cost implication: Thinking tokens are billed at the same rate as output tokens. Choosing low effort reduces the number of reasoning tokens the model generates, which directly reduces your cost per request. For classification or extraction tasks, low can cut thinking token usage by 60–80% compared to max.

Responses API: What It Means for Codex and Agent Workflows

DeepSeek's Responses API is a drop-in implementation of the OpenAI Responses API format. Supported features:

  • Streaming with semantic SSE events — response.output_text.delta, response.reasoning_text.delta, and typed tool call events
  • Function calling and web search via tools parameter
  • Reasoning effort via the reasoning.effort field
  • instructions parameter — inserted as the first system message
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_DEEPSEEK_KEY",
    base_url="https://api.deepseek.com",
)

response = client.responses.create(
    model="deepseek-v4-pro",
    instructions="You are a helpful coding assistant.",
    input="Refactor this function to use async/await.",
)

print(response.output_text)

Notable limitations compared to the OpenAI Responses API: previous_response_id, conversation, store, background, and metadata are not supported. DeepSeek's implementation is stateless — each request is independent. Unsupported parameters are silently ignored, so existing clients do not break.

Codex Integration

DeepSeek provides a one-click Codex setup with a model configuration that includes multi-agent v2 support, apply_patch tool type set to freeform, and a 1M token context window. The configuration supports both V4-Flash and V4-Pro models.

Routing DeepSeek V4-Pro Through TheRouter

TheRouter routes OpenAI-compatible requests through configured providers, which means you can use DeepSeek V4-Pro as a primary model and configure fallback routing to other providers when DeepSeek is unavailable.

A basic configuration that routes deepseek-v4-pro requests to DeepSeek directly:

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_THEROUTER_KEY",
    base_url="https://therouter.ai/api/v1",
)

response = client.chat.completions.create(
    model="deepseek/deepseek-v4-pro",
    messages=[{"role": "user", "content": "Explain peak pricing"}],
)

Using DashScope as a Fallback Path

V4-Pro is also available on DashScope (Alibaba Cloud Bailian) under the model ID deepseek-v4-pro-0813. DashScope uses a different pricing structure (RMB-denominated, no peak/off-peak split as of this writing), which can serve as a cost-stabilizing fallback when DeepSeek's peak rates apply.

Scheduling Workloads for Off-Peak Savings

The peak/off-peak model creates a straightforward optimization strategy: defer deferrable work to off-peak windows.

Workloads that can be deferred:

  • Batch evaluation runs
  • Dataset labeling and annotation
  • Document summarization pipelines
  • Code review passes on non-urgent PRs
  • Embedding generation for RAG indexes

Workloads that cannot be deferred:

  • Interactive chat sessions
  • Real-time coding assistant responses
  • Latency-sensitive API endpoints

Implementation Approaches

1. Cron-scheduled batch jobs. Run your batch processing during off-peak hours. For a Beijing-based team, that means scheduling jobs outside 09:00–12:00 and 14:00–18:00 CST.

2. Queue-based deferral. Push non-urgent requests to a queue (SQS, RabbitMQ, Redis) and drain the queue during off-peak windows. This works well for document processing pipelines.

3. Provider-level Batch API. DeepSeek does not currently offer a Batch API equivalent to OpenAI's, but the off-peak pricing effectively serves the same purpose — you get a 50% discount for processing during low-demand windows, compared to OpenAI's 50% discount on Batch API calls.

Common Errors and Fixes

ErrorCauseFix
400: reasoning_content must participate in contextTool call in thinking mode — you did not pass back reasoning_contentInclude reasoning_content from the assistant message in subsequent turns when tool calls are involved
400: request exceeds context windowInput exceeds 1M tokensTruncate input or split into multiple requests
Rate limit exceededExceeded 500 concurrent requests (V4-Pro) or 2,500 (V4-Flash)Add backoff logic, or route overflow to a secondary provider

Production Checklist

Before going live with V4-Pro in a production environment:

  • Confirm you are using the deepseek-v4-pro model name (not legacy deepseek-chat or deepseek-reasoner)
  • Set reasoning_effort explicitly per use case — do not accept the high default for simple tasks
  • Monitor peak vs. off-peak spend — if more than 60% of your token volume lands in peak windows, evaluate whether you can shift batch workloads
  • Configure fallback routing to at least one alternative provider (V4-Flash, DashScope-hosted V4-Pro, or a different model family entirely)
  • Test Responses API compatibility if you plan to use Codex or other Responses API consumers
  • Set up alerting on the 500-request concurrency limit for V4-Pro

Further Reading

Models covered in this article

Help & contact