DeepSeek V4 Pro Permanent Price Cut: 75% Discount Stays — $0.87/M Output Reshapes Routing Tiers
DeepSeek V4 Pro's 75% price cut is now permanent — $0.87/M output tokens with 80%+ SWE-bench scores. Here's why it shifts V4 Pro from premium fallback to primary reasoning tier and what that means for your routing cost structure.
Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

The AI price war just shifted from promotional tactics to permanent strategy. DeepSeek confirmed on May 23, 2026 that the 75% discount on V4 Pro—originally set to expire May 31—will not revert. What was launched in April at $1.74/M input and $3.48/M output now permanently sits at $0.435/M input and $0.87/M output. That is not a rounding difference. It is a category change.
This matters less as a news item and more as a routing policy trigger: every team that configured V4 Pro as a "premium fallback for complex tasks" needs to revisit whether it should now be the primary tier.
What changed
DeepSeek launched V4 Pro on April 24, 2026 alongside V4 Flash, positioning them as a split-tier lineup:
| Model | Input (cache miss) | Output | Context | Thinking |
|---|---|---|---|---|
| deepseek-v4-flash | $0.14/M | $0.28/M | 1M | Both modes |
| deepseek-v4-pro | $0.435/M | $0.87/M | 1M | Both modes |
The launch price for V4 Pro was $1.74/M input and $3.48/M output. A promotional 75% discount ran from launch through May 31, but the company statement confirmed on May 23 that these discounted rates are now the permanent list prices. Cache hit pricing was also reduced to 1/10 of launch rates across both models from April 26.
Concurrency limits remain split: V4 Flash supports 2,500 concurrent requests; V4 Pro allows 500.
Why it matters for AI engineering teams
The cost comparison with comparable frontier models has shifted permanently. At $0.87/M output tokens, V4 Pro is now meaningfully cheaper than:
- OpenAI o4-mini: $4.40/M output (5× more expensive)
- Claude Sonnet 4: $15/M output (17× more expensive)
- GPT-4.1: $8/M output (9× more expensive)
DeepSeek V4 Pro scores 80.6% on SWE-bench Verified—comparable to or exceeding several models that cost an order of magnitude more. It ships with 1M token context, tool calling, JSON output, both thinking and non-thinking modes, and dual OpenAI/Anthropic API format compatibility.
The original framing was: use V4 Pro only when V4 Flash cannot handle the task. The permanent pricing changes that framing: V4 Pro is now priced as a default reasoning model, not a premium exception.
The "fallback economics" calculation changes. Teams that built routing policies around "fall back to cheaper domestic model on V4 Pro failures" may now be routing the math backwards. At these prices, the fallback tail cost is small enough that V4 Pro as the default primary tier—with V4 Flash as the fast, cheap execution tier—may be more cost-effective than routing V4 Pro as the fallback.
China-to-global routing tension sharpens. For teams using multiple providers across regions, V4 Pro's permanent pricing makes it a stronger anchor for reasoning-heavy workloads even against newer global models. The DeepSeek API also now exposes both OpenAI-format and Anthropic-format base URLs, making multi-provider routing simpler from a protocol perspective.
The router/operator angle
The permanent price change creates three immediate routing policy decisions:
1. Revise your model tier assignment. If V4 Pro was configured as "tier 2 fallback," evaluate whether it should be "tier 1 primary for reasoning" with V4 Flash as the "tier 0 fast/cheap" execution model. The $0.87/M output price makes V4 Pro competitive with flash-tier models from many other providers.
2. Revisit thinking-mode economics. Both V4 Pro and V4 Flash default to thinking mode enabled. Thinking mode increases output token count—which directly multiplies your per-query cost. At $0.87/M output with thinking on, a 4,000-token thinking chain costs $0.0035. At $3.48 (old price), the same chain cost $0.014. The per-query math is 4× better; the thinking-mode opt-in decision is more forgiving than it was.
3. Audit your fallback chain for reverse-routing risk. A common pattern: route to cheap domestic model first, fall back to expensive global model on failure. If you set that fallback as Claude Sonnet or GPT-4.1, and your primary is now V4 Pro at $0.87/M output, your fallback on failure is suddenly 10-17× more expensive per token. Make sure your fallback chains reflect current pricing, not the pricing at the time you wrote the config.
Routing tier framework for 2026-05 pricing landscape:
| Tier | Model | Output $/M | Use case |
|---|---|---|---|
| Fast/economy | deepseek-v4-flash | $0.28 | Autocomplete, quick evals, retries |
| Primary reasoning | deepseek-v4-pro | $0.87 | Code, long-context, agent tasks |
| Frontier reasoning | o4-mini | $4.40 | Tasks requiring strongest available model |
| Frontier completion | Claude Sonnet 4 | $15 | Tasks needing specific Anthropic capabilities |
Concurrency ceiling risk. V4 Pro caps at 500 concurrent requests. Teams with bursting workloads should build concurrency-based overflow routing: when V4 Pro is saturated, route overflow to V4 Flash or a secondary provider rather than failing or queuing indefinitely. This is the same pattern as capacity-aware routing for any provider with concurrency limits.
What TheRouter users should watch or try
If you route to DeepSeek through TheRouter, the pricing change is transparent—your provider config does not need updating if you already migrated from deepseek-chat to deepseek-v4-pro. The price you see in billing will simply be lower.
The routing policy updates that do require action are logical, not mechanical:
- Review whether your provider ordering still reflects the "cheapest-first" or "quality-first" intent you designed for. V4 Pro's cost position has moved, and the order may now be wrong.
- If you configured V4 Pro with a cost-based routing guard (e.g., "use Pro only for prompts over N tokens"), revisit whether that threshold still makes sense at $0.87/M output.
- For Claude Code or OpenCode users: V4 Pro via
deepseek.com/anthropicendpoint now hits a better price-performance point for long-horizon agentic tasks. Teams that set up direct Anthropic-format routing to DeepSeek as described in the V4 migration guide are getting a permanent pricing benefit with no config change required.
See the SiliconFlow provider guide for an example of configuring a Chinese AI provider through TheRouter's routing layer—the same cost-control principles apply to DeepSeek V4 Pro today.
The broader trend: permanent price compression from frontier-class Chinese models is now a structural feature of the AI provider landscape, not a promotional window. Routing policies need to be designed to take advantage of this, not just treat it as temporary arbitrage.
Models covered in this article

GLM-5.2 Fast Mode Gets a 20% Price Cut on DashScope: What Alibaba's Move Means for Your Routing Cost Model
Alibaba Cloud Bailian cut the GLM-5.2 Fast mode token price by 20% on July 15. For operators routing cost-sensitive workloads through DashScope's OpenAI-compatible endpoint, the math just changed.

DeepSeek V4 Pro Survives Its Own Deadline: What the Reversal Means for Your Routing Policy
DeepSeek announced on September 10 that V4 Pro would be retired today at 04:00 UTC. Instead, they reversed course in response to user demand, keeping V4 Pro live at unchanged pricing. Here is what the two-model landscape now looks like and which workloads belong on which.

DeepSeek V4.1-Flash Lands With Native Vision — and Kills deepseek-v4-pro in Four Days
DeepSeek V4.1-Flash ships today under a new API name with native multimodal. On September 14, every deepseek-v4-pro request silently reroutes to V4.1-Flash at Flash pricing. Here is what the capability delta means and what to pin before the deadline.