DeepSeek V4-Pro-0813 and the August 16 Pricing Cliff: What Every Gateway Operator Must Do in the Next 48 Hours
DeepSeek's V4-Pro alias just upgraded to a new GA model with top agent benchmarks, and peak/off-peak pricing goes live August 16. The cost math for production routing teams has fundamentally changed.

Two things happened this week that together change the arithmetic for every team running production traffic through DeepSeek's API. First, DeepSeek promoted V4-Pro-0813 to GA — the deepseek-v4-pro alias now points to it. Second, the peak/off-peak pricing structure that circulated as a rumor in early July has been officially confirmed, with exact figures and deadline: 16:00 UTC on August 16, 2026 — less than 48 hours from now.
If your routing policy, cron jobs, or batch pipelines were configured based on flat DeepSeek pricing, they need to change today.
The V4-Pro-0813 Model Upgrade
The GA release consolidates improvements across agent benchmarks that were previewed in earlier testing. The verified numbers from DeepSeek's changelog: HLE (without/with tools) 42.7/60.0, TerminalBench 2.1 at 87.9, NL2Repo 61.5, Cybergym 83.3, DeepSWE 62.7, Toolathlon-Verified 74.1, DSBench-FullStack 71.1, DSBench-Hard 67.2.
The more operationally significant change is the thinking effort control. The model now supports three discrete levels via the reasoning_effort parameter: "low", "high", and "max". The older binary approach — setting thinking: {type: "enabled"} — still works, but reasoning_effort is now the recommended control. For routing teams, the practical mapping is:
reasoning_effort: "low"— cost-sensitive simple tasks; fastest, cheapestreasoning_effort: "high"— recommended default for daily agent tasks; good balancereasoning_effort: "max"— complex multi-step reasoning, planning, and hard coding problems
The model also gains native Responses API support, specifically adapted for Codex workflows. If you're running Codex integrations against deepseek-v4-pro, this is a one-config update.
The Pricing Cliff — Exact Numbers, Hard Deadline
In early July, a report based on a TechNode rumor suggested DeepSeek would introduce peak-hour pricing with vague figures. That story — published here in July — was based on secondhand information. Today's confirmed prices from DeepSeek's official pricing page are different from what was rumored, and the timing is now exact.
Current prices (in effect until 16:00 UTC August 16):
| Model | Cache miss input | Output |
|---|---|---|
| deepseek-v4-flash | $0.14/M | $0.28/M |
| deepseek-v4-pro | $0.435/M | $0.87/M |
Prices after 16:00 UTC August 16:
| Model | Window | Cache hit | Cache miss input | Output |
|---|---|---|---|---|
| deepseek-v4-flash | Off-peak | $0.007/M | $0.22/M | $0.66/M |
| deepseek-v4-flash | Peak | $0.014/M | $0.44/M | $1.32/M |
| deepseek-v4-pro | Off-peak | $0.022/M | $0.66/M | $1.98/M |
| deepseek-v4-pro | Peak | $0.044/M | $1.32/M | $3.96/M |
Peak hours are 01:00–04:00 UTC and 06:00–10:00 UTC. Everything else is off-peak.
The multiplier between peak and off-peak is exactly 2×. The multiplier between today's flat rate and tomorrow's peak rate for deepseek-v4-pro output is 4.55×.
What This Means for Gateway Operators
The peak windows cover 01:00–04:00 UTC (US West Coast late evening, Asia early morning) and 06:00–10:00 UTC (Western Europe working hours, US East Coast morning). For most production teams, this is squarely in primary working hours.
An operator routing 10M output tokens/day to deepseek-v4-pro entirely during peak hours will pay $39.60/day starting August 16. The same workload shifted entirely to off-peak hours costs $19.80/day. At scale, a team that schedules batch processing intelligently versus one that doesn't will see a 2× cost difference on the same model for the same work.
To put it sharply: routing to deepseek-v4-pro during the 06:00–10:00 UTC window at $3.96/M output is more expensive than routing to most flat-rate alternatives during those hours. Time-of-day is now a first-class routing dimension for DeepSeek traffic.
This is meaningfully different from OpenAI's Fast Mode (service_tier-based), which reduces latency and improves throughput but does not introduce time-of-day cost variation. Teams evaluating provider mix can use OpenAI Fast Mode for cost-predictable peak-hour traffic while shifting bulk async work to DeepSeek off-peak windows — a strategy that was not available before this pricing structure was confirmed.
What to Audit Before August 16
Scheduled jobs: Any cron job or pipeline that fires during 01:00–04:00 UTC or 06:00–10:00 UTC will pay peak rates on DeepSeek calls. Move heavy batch processing to UTC 10:00–23:59 or the 04:00–06:00 gap.
Reasoning effort settings: If you've been using thinking: {type: "enabled"} without a budget, you are now at risk of unintentionally hitting max-tier compute under the new model. Set reasoning_effort explicitly: start with "high" for most agent tasks, and reserve "max" for tasks where quality justifies the cost.
Codex integrations: If you use the native Responses API path with DeepSeek, update your config to use the deepseek-v4-pro alias with the Responses API as documented — DeepSeek's changelog notes it is now specifically adapted for Codex workflows.
Cost forecasts: Models built on flat per-token rates are now inaccurate for DeepSeek. Adjust any budget projections that use DeepSeek V4 Pro output pricing — the right number depends entirely on when your traffic actually runs.
The pricing change is substantial. The model upgrade is real. Both take effect around the same time, and they need to be addressed together in routing config, scheduling policy, and cost accounting before 16:00 UTC tomorrow.

DeepSeek V4 Pro Survives Its Own Deadline: What the Reversal Means for Your Routing Policy
DeepSeek announced on September 10 that V4 Pro would be retired today at 04:00 UTC. Instead, they reversed course in response to user demand, keeping V4 Pro live at unchanged pricing. Here is what the two-model landscape now looks like and which workloads belong on which.

DeepSeek V4.1-Flash Lands With Native Vision — and Kills deepseek-v4-pro in Four Days
DeepSeek V4.1-Flash ships today under a new API name with native multimodal. On September 14, every deepseek-v4-pro request silently reroutes to V4.1-Flash at Flash pricing. Here is what the capability delta means and what to pin before the deadline.

DeepSeek V4 Peak-Hour Pricing Makes Time-of-Day a Routing Dimension Every AI Gateway Must Support
DeepSeek V4 launches in mid-July with 2× peak-hour API pricing — the first surge pricing in AI APIs. Every routing layer now needs time-zone-aware cost accounting and fallback logic.