← All articles

DeepSeek V4-Pro vs V4-Flash (August 2026): API Pricing, Benchmarks, and Routing Strategies

A head-to-head comparison of DeepSeek V4-Pro-0813 and V4-Flash-0731 covering peak/off-peak pricing, independent benchmarks, reasoning effort control, and concrete routing strategies for OpenAI-compatible API operators.

· updated 2026-08-18· TheRouter

DeepSeek now ships two production models behind its OpenAI-compatible endpoint: V4-Pro (version 0813, GA since August 13) and V4-Flash (version 0731, GA since July 31). Both share a 1M-token context window, 384K maximum output, hybrid thinking mode, tool calling, JSON output, the Responses API, and an Anthropic-format endpoint. On paper they look like the same model at two price points. In practice, the gap is narrower than those names suggest — and the routing decision comes down to workload sensitivity, not a blanket quality difference.

This comparison lays out the current pricing (including the new peak/off-peak rates effective August 16), the independent benchmark evidence, and a concrete routing framework for operators running both models behind a single base_url.

OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.

Pricing: Peak, Off-Peak, and the 3× Multiplier

DeepSeek introduced time-of-day pricing on August 16, 2026. Peak hours are 01:00–04:00 UTC and 06:00–10:00 UTC (roughly 09:00–12:00 and 14:00–18:00 Beijing time). Off-peak rates are exactly half of peak rates.

All prices per 1M tokens, in USD:

Price ComponentV4-Flash (off-peak)V4-Flash (peak)V4-Pro (off-peak)V4-Pro (peak)Pro/Flash Ratio
Cache-hit input$0.007$0.014$0.022$0.044~3.1×
Cache-miss input$0.22$0.44$0.66$1.323×
Output$0.66$1.32$1.98$3.963×
Concurrency limit2,5002,5005005000.2×

Three things stand out:

  1. Pro costs exactly 3× Flash across cache-miss input and output at both peak and off-peak rates. The ratio is consistent — no surprise multipliers on output tokens.

  2. Off-peak is half of peak for both models. An operator who can shift batch workloads to off-peak hours (10:00–01:00 UTC, 04:00–06:00 UTC) pays the equivalent of the old pre-August-16 rates on Flash and significantly less on Pro.

  3. Flash has 5× the concurrency. At 2,500 concurrent requests vs Pro's 500, Flash is the only viable choice for high-throughput fan-out workloads like bulk classification or parallel sub-agent execution.

When comparing API pricing across providers, always normalize to USD per million tokens and split input from output. Most providers price output tokens 2–5× higher than input tokens, so a workload heavy on completion length looks very different from a retrieval-heavy workload at the same nominal "price per million."

  • Use one currency (USD) — convert at publish date and cite the rate.
  • Split input/output — never quote a single blended number.
  • Cite each row to the provider's own pricing page with retrieval date.
  • Note context-window tiers — long-context pricing often steps higher.

What These Prices Mean in Practice

For a typical chat turn (1K fresh input + 500 output tokens) at off-peak rates:

  • V4-Flash: ~$0.00055
  • V4-Pro: ~$0.00165

For a cache-heavy agent loop (200K cached + 20K fresh input + 10K output) at off-peak rates:

  • V4-Flash: ~$0.011
  • V4-Pro: ~$0.033

The 3× multiplier holds across workload shapes. Pro never becomes disproportionately expensive for long contexts or heavy output — it is a flat premium.

Independent Benchmarks: Closer Than the Names Suggest

The "Pro" and "Flash" labels imply a large capability gap. Independent testing tells a different story.

Data from Artificial Analysis (checked August 13, 2026):

BenchmarkV4-Pro 0813V4-Flash 0731Gap
Intelligence Index5352Pro +1
Agentic Index49.648.4Pro +1.2
Terminal-Bench v2.178.65%78.65%Tied
Long-context reasoning75.33%74.33%Pro +1
GPQA Diamond92.83%90.81%Pro +2
SciCode49.19%49.88%Flash +0.7
Output speed83.2 tok/s122.2 tok/sFlash 47% faster

DeepSeek's own benchmarks (from the V4 technical report on Hugging Face) show a wider spread on knowledge and agentic tasks:

BenchmarkV4-Pro 0813V4-Flash 0731Gap
SimpleQA57.9%34.1%Pro +23.8
BrowseComp83.4%73.2%Pro +10.2
Terminal-Bench 2.067.9%56.9%Pro +11
SWE-bench Pro55.4%52.6%Pro +2.8
HLE42.7%34.8%Pro +7.9

The independent and vendor data paint different pictures. On independent evaluations, Pro and Flash land in the same broad capability tier — the gap is single-digit on most metrics. On DeepSeek's own benchmarks (Terminal-Bench 2.0, SimpleQA, BrowseComp), Pro pulls ahead more substantially.

Three conclusions survive both datasets:

  1. V4-Pro 0813 is a significant upgrade over the earlier Pro preview. Artificial Analysis recorded jumps from ~45.3 to 53 on the Intelligence Index and from ~37.8 to 49.6 on the Agentic Index.

  2. For routine coding and agent tasks, Flash matches or nearly matches Pro on independent evaluations. Terminal-Bench v2.1 is tied. SWE-bench Pro shows a 2.8-point gap.

  3. Pro's advantage concentrates on knowledge-heavy and multi-step agentic tasks — SimpleQA, BrowseComp, HLE — where the cost of a wrong first pass is high.

Feature Parity

Both models share the same API surface:

FeatureV4-ProV4-Flash
Context window1M1M
Max output384K384K
Thinking modeDefault on, effort controlDefault on, effort control
Reasoning effort levelslow / high / maxlow / high / max
Tool callingYesYes
JSON outputYesYes
Responses APIYesYes
Anthropic API formatYesYes
FIM completion (beta)Non-thinking onlyNon-thinking only
Chat prefix completion (beta)YesYes

The reasoning effort control is identical on both models. Setting reasoning_effort: "low" maps to low actual effort; "high" maps to high; "max" maps to max. The intermediate values ("medium", "xhigh") both map to high. This means you can use effort control as a per-request cost lever without model switching — drop to "low" for simple classification, use "high" for production agent loops, escalate to "max" for complex multi-step reasoning.

Routing Strategy: When to Use Which

The routing decision between V4-Pro and V4-Flash is not "use Pro for hard tasks and Flash for easy ones." It is "use Flash as the default and escalate to Pro when the cost of a wrong first answer exceeds the 3× token premium."

Start with Flash

Flash should be your default route for:

  • High-volume code generation — retryable, measurable, and 47% faster
  • Classification, extraction, and summarization — tasks where quality differences are small and throughput matters
  • Sub-agent workers — the 2,500 concurrency limit supports parallel fan-out
  • Exploratory prototyping — establish a cost baseline before paying for Pro

Escalate to Pro When

Pro earns its 3× premium when:

  • Architecture planning and cross-system debugging — errors propagate and create expensive rework
  • Knowledge-intensive queries — Pro's SimpleQA gap (57.9% vs 34.1%) suggests stronger factual grounding
  • High-stakes analysis — when the cost of human review on a wrong answer exceeds the API cost difference
  • Multi-step agentic tasks with branching — BrowseComp and Terminal-Bench 2.0 gaps suggest Pro handles complex agent loops better

The Planner-Worker Pattern

A common production pattern: use Pro as the planner (architecture decisions, task decomposition) and Flash as the worker (code generation, test execution, file operations). This concentrates the premium model where judgment matters and lets Flash handle the volume.

# Example: route by task type through TheRouter
from openai import OpenAI

client = OpenAI(
    base_url="https://therouter.ai/v1",
    api_key="your-therouter-key"
)

# Planning tasks → V4-Pro
plan = client.chat.completions.create(
    model="deepseek/deepseek-v4-pro",
    messages=[{"role": "user", "content": "Design the migration plan..."}],
    reasoning_effort="high"
)

# Execution tasks → V4-Flash
for task in plan_tasks:
    result = client.chat.completions.create(
        model="deepseek/deepseek-v4-flash",
        messages=[{"role": "user", "content": task}],
        reasoning_effort="low"
    )

Optimizing with Peak/Off-Peak Scheduling

Since off-peak rates are 50% of peak rates, operators running batch or non-latency-sensitive workloads should schedule them during off-peak hours (10:00–01:00 UTC, 04:00–06:00 UTC). Combined with Flash as the default model, this yields up to 6× savings compared to Pro at peak rates.

For latency-sensitive production traffic, the peak/off-peak distinction is less actionable — you serve requests when they arrive. But batch evaluation, regression testing, and data pipeline jobs can all shift to off-peak windows.

Routing Through TheRouter

Both models are available in TheRouter's model catalog as deepseek/deepseek-v4-pro and deepseek/deepseek-v4-flash. You can set up model fallbacks to automatically route between them:

  • Primary: deepseek/deepseek-v4-flash for cost-efficient default routing
  • Fallback: deepseek/deepseek-v4-pro when Flash returns errors or hits concurrency limits
  • Or: route by task metadata — send planning prompts to Pro, execution prompts to Flash

The OpenAI-compatible base_url means switching models requires changing only the model parameter. No SDK changes, no endpoint changes, no authentication changes.

When V4-Pro Is Not Worth It

Pro is not worth the premium when:

  • You have no task-level evaluation data yet — establish a Flash baseline first
  • The task is retryable with cheap verification (unit tests, format checks)
  • Throughput matters more than per-request quality — Flash's 5× concurrency and 47% speed advantage dominate
  • You are running during peak hours and cost is the binding constraint — Flash at peak ($1.32/1M output) is cheaper than Pro at off-peak ($1.98/1M output)

Bottom Line

V4-Pro 0813 is a real upgrade over the earlier preview and pulls ahead on knowledge-heavy and complex agentic benchmarks. V4-Flash 0731 matches or nearly matches Pro on independent coding and reasoning evaluations while running 47% faster, supporting 5× more concurrent requests, and costing exactly one-third the price.

For most production workloads, Flash is the correct default. Pro earns its place in specific high-stakes scenarios where the cost of a wrong first answer exceeds the 3× token premium. The planner-worker pattern — Pro for planning, Flash for execution — captures the best of both without overspending.

With peak/off-peak pricing now live, operators who can schedule batch work during off-peak hours compound their savings further. Route both models through a single OpenAI-compatible gateway and let your workload data, not model names, drive the routing decision.


Sources:

Models covered in this article

Help & contact