DeepSeek V4 Pro Permanent Price Cut: 75% Discount Stays — $0.87/M Output Reshapes Routing Tiers

DeepSeek V4 Pro's 75% price cut is now permanent — $0.87/M output tokens with 80%+ SWE-bench scores. Here's why it shifts V4 Pro from premium fallback to primary reasoning tier and what that means for your routing cost structure.

Published via DeepSeek

Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

DeepSeek V4 Pro permanent price cut: routing tier decision framework for AI engineering teams

The AI price war just shifted from promotional tactics to permanent strategy. DeepSeek confirmed on May 23, 2026 that the 75% discount on V4 Pro—originally set to expire May 31—will not revert. What was launched in April at $1.74/M input and $3.48/M output now permanently sits at $0.435/M input and $0.87/M output. That is not a rounding difference. It is a category change.

This matters less as a news item and more as a routing policy trigger: every team that configured V4 Pro as a "premium fallback for complex tasks" needs to revisit whether it should now be the primary tier.

What changed

DeepSeek launched V4 Pro on April 24, 2026 alongside V4 Flash, positioning them as a split-tier lineup:

ModelInput (cache miss)OutputContextThinking
deepseek-v4-flash$0.14/M$0.28/M1MBoth modes
deepseek-v4-pro$0.435/M$0.87/M1MBoth modes

The launch price for V4 Pro was $1.74/M input and $3.48/M output. A promotional 75% discount ran from launch through May 31, but the company statement confirmed on May 23 that these discounted rates are now the permanent list prices. Cache hit pricing was also reduced to 1/10 of launch rates across both models from April 26.

Concurrency limits remain split: V4 Flash supports 2,500 concurrent requests; V4 Pro allows 500.

Why it matters for AI engineering teams

The cost comparison with comparable frontier models has shifted permanently. At $0.87/M output tokens, V4 Pro is now meaningfully cheaper than:

  • OpenAI o4-mini: $4.40/M output (5× more expensive)
  • Claude Sonnet 4: $15/M output (17× more expensive)
  • GPT-4.1: $8/M output (9× more expensive)

DeepSeek V4 Pro scores 80.6% on SWE-bench Verified—comparable to or exceeding several models that cost an order of magnitude more. It ships with 1M token context, tool calling, JSON output, both thinking and non-thinking modes, and dual OpenAI/Anthropic API format compatibility.

The original framing was: use V4 Pro only when V4 Flash cannot handle the task. The permanent pricing changes that framing: V4 Pro is now priced as a default reasoning model, not a premium exception.

The "fallback economics" calculation changes. Teams that built routing policies around "fall back to cheaper domestic model on V4 Pro failures" may now be routing the math backwards. At these prices, the fallback tail cost is small enough that V4 Pro as the default primary tier—with V4 Flash as the fast, cheap execution tier—may be more cost-effective than routing V4 Pro as the fallback.

China-to-global routing tension sharpens. For teams using multiple providers across regions, V4 Pro's permanent pricing makes it a stronger anchor for reasoning-heavy workloads even against newer global models. The DeepSeek API also now exposes both OpenAI-format and Anthropic-format base URLs, making multi-provider routing simpler from a protocol perspective.

The router/operator angle

The permanent price change creates three immediate routing policy decisions:

1. Revise your model tier assignment. If V4 Pro was configured as "tier 2 fallback," evaluate whether it should be "tier 1 primary for reasoning" with V4 Flash as the "tier 0 fast/cheap" execution model. The $0.87/M output price makes V4 Pro competitive with flash-tier models from many other providers.

2. Revisit thinking-mode economics. Both V4 Pro and V4 Flash default to thinking mode enabled. Thinking mode increases output token count—which directly multiplies your per-query cost. At $0.87/M output with thinking on, a 4,000-token thinking chain costs $0.0035. At $3.48 (old price), the same chain cost $0.014. The per-query math is 4× better; the thinking-mode opt-in decision is more forgiving than it was.

3. Audit your fallback chain for reverse-routing risk. A common pattern: route to cheap domestic model first, fall back to expensive global model on failure. If you set that fallback as Claude Sonnet or GPT-4.1, and your primary is now V4 Pro at $0.87/M output, your fallback on failure is suddenly 10-17× more expensive per token. Make sure your fallback chains reflect current pricing, not the pricing at the time you wrote the config.

Routing tier framework for 2026-05 pricing landscape:

TierModelOutput $/MUse case
Fast/economydeepseek-v4-flash$0.28Autocomplete, quick evals, retries
Primary reasoningdeepseek-v4-pro$0.87Code, long-context, agent tasks
Frontier reasoningo4-mini$4.40Tasks requiring strongest available model
Frontier completionClaude Sonnet 4$15Tasks needing specific Anthropic capabilities

Concurrency ceiling risk. V4 Pro caps at 500 concurrent requests. Teams with bursting workloads should build concurrency-based overflow routing: when V4 Pro is saturated, route overflow to V4 Flash or a secondary provider rather than failing or queuing indefinitely. This is the same pattern as capacity-aware routing for any provider with concurrency limits.

What TheRouter users should watch or try

If you route to DeepSeek through TheRouter, the pricing change is transparent—your provider config does not need updating if you already migrated from deepseek-chat to deepseek-v4-pro. The price you see in billing will simply be lower.

The routing policy updates that do require action are logical, not mechanical:

  • Review whether your provider ordering still reflects the "cheapest-first" or "quality-first" intent you designed for. V4 Pro's cost position has moved, and the order may now be wrong.
  • If you configured V4 Pro with a cost-based routing guard (e.g., "use Pro only for prompts over N tokens"), revisit whether that threshold still makes sense at $0.87/M output.
  • For Claude Code or OpenCode users: V4 Pro via deepseek.com/anthropic endpoint now hits a better price-performance point for long-horizon agentic tasks. Teams that set up direct Anthropic-format routing to DeepSeek as described in the V4 migration guide are getting a permanent pricing benefit with no config change required.

See the SiliconFlow provider guide for an example of configuring a Chinese AI provider through TheRouter's routing layer—the same cost-control principles apply to DeepSeek V4 Pro today.

The broader trend: permanent price compression from frontier-class Chinese models is now a structural feature of the AI provider landscape, not a promotional window. Routing policies need to be designed to take advantage of this, not just treat it as temporary arbitrage.

Models covered in this article

Help & contact