DeepSeek V4 Pro Retires September 14: How to Migrate to V4.1 Flash
DeepSeek V4 Pro will be retired on September 14, 2026. All V4 Pro requests will route to V4.1 Flash at Flash pricing. We walk through the model-name changes, code updates, pricing impact, and a rollback plan for operators who rely on V4 Pro today.
DeepSeek V4 Pro retires on September 14, 2026, at 12:00 Beijing Time (04:00 UTC). From that moment, every request sent to deepseek-v4-pro will be silently rerouted to V4.1 Flash and billed at the Flash price. This is not a gradual sunset with months of overlap. It is a hard cutover in four days.
The V4 Flash model has already gone through the same transition. The names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted by the API, but both now resolve to V4.1 Flash under the canonical name deepseek-flash. DeepSeek's stated reason for the consolidation is that V4.1 Flash "has comprehensively surpassed V4 Pro in performance, cost, speed, and total time" (source, retrieved 2026-09-10).
We use DeepSeek as a routing target in TheRouter, so this affects our stack directly. This guide covers what changed, what you need to do, and where the risks are.
Timeline Summary
| Date | Event | Impact |
|---|---|---|
| April 24, 2026 | V4 Preview launch | deepseek-v4-pro and deepseek-v4-flash introduced |
| Before Sep 10 | V4 Flash retired | deepseek-v4-flash and deepseek-v4-flash-vision-exp now route to V4.1 Flash under the name deepseek-flash |
| Sep 14, 2026 | V4 Pro retires | deepseek-v4-pro reroutes to V4.1 Flash, billed at Flash price |
| Not announced | V4.1 Pro release | DeepSeek has announced a future V4.1 Pro but no date or pricing |
Model Name Changes
This is the part most likely to break your config. DeepSeek consolidated three model names into one.
Before (pre-September 2026):
deepseek-v4-flash → DeepSeek-V4-Flash
deepseek-v4-flash-vision-exp → DeepSeek-V4-Flash (vision)
deepseek-v4-pro → DeepSeek-V4-Pro-0813
After (September 14 onward):
deepseek-flash → DeepSeek-V4.1-Flash (canonical name)
deepseek-v4-flash → DeepSeek-V4.1-Flash (legacy alias, still accepted)
deepseek-v4-flash-vision-exp → DeepSeek-V4.1-Flash (legacy alias, still accepted)
deepseek-v4-pro → DeepSeek-V4.1-Flash (rerouted, billed at Flash price)
The safest move is to update your model parameter to deepseek-flash now. The legacy names continue to work, but they mask which model you are actually hitting.
Code Update
If you are using the OpenAI-compatible SDK, the change is a single line.
Before:
from openai import OpenAI
client = OpenAI(
api_key="YOUR_DEEPSEEK_KEY",
base_url="https://api.deepseek.com",
)
response = client.chat.completions.create(
model="deepseek-v4-pro", # retiring Sep 14
messages=[{"role": "user", "content": "Explain async generators in Python."}],
)
After:
from openai import OpenAI
client = OpenAI(
api_key="YOUR_DEEPSEEK_KEY",
base_url="https://api.deepseek.com",
)
response = client.chat.completions.create(
model="deepseek-flash", # V4.1 Flash — canonical name
messages=[{"role": "user", "content": "Explain async generators in Python."}],
)
The same applies if you use the Anthropic-compatible endpoint (base_url="https://api.deepseek.com/anthropic"). Update the model parameter from deepseek-v4-pro to deepseek-flash.
Pricing Impact
V4.1 Flash is cheaper than V4 Pro across the board. After September 14, even if you keep sending deepseek-v4-pro, you will be billed at the Flash rate.
| Metric | V4 Pro (current) | V4.1 Flash (after Sep 14) | Change |
|---|---|---|---|
| Input (cache miss, off-peak) | $0.66 / 1M tokens | $0.15 / 1M tokens | -77% |
| Input (cache miss, peak) | $1.32 / 1M tokens | $0.30 / 1M tokens | -77% |
| Input (cache hit, off-peak) | $0.022 / 1M tokens | $0.003 / 1M tokens | -86% |
| Input (cache hit, peak) | $0.044 / 1M tokens | $0.006 / 1M tokens | -86% |
| Output (off-peak) | $1.98 / 1M tokens | $0.60 / 1M tokens | -70% |
| Output (peak) | $3.96 / 1M tokens | $1.20 / 1M tokens | -70% |
| Concurrency limit | 500 | 2,500 | +5x |
Pricing data sourced from the DeepSeek pricing page (retrieved 2026-09-10). Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday.
The cost reduction is real. The question is whether the quality trade-off matters for your workload.
What V4.1 Flash Changes in Practice
V4.1 Flash supports the same feature set as V4 Pro with two exceptions.
Gained: Vision support. deepseek-flash accepts image inputs. V4 Pro did not support vision.
Same: Thinking mode (enabled by default, effort levels low/high/max), tool calls, JSON output, Responses API, FIM completion (non-thinking only), Anthropic API format, 1M context window, 384K max output.
Risk area: DeepSeek claims V4.1 Flash surpasses V4 Pro on all axes. If your workload depends on V4 Pro's specific behavior for complex reasoning or agentic coding, test before September 14. DeepSeek has not published independent benchmark comparisons between V4 Pro and V4.1 Flash.
Rate Limit Changes
The concurrency limit jumps from 500 (V4 Pro) to 2,500 (V4.1 Flash). This is a pure upgrade for throughput-sensitive workloads. No action needed — the higher limit applies automatically once your requests route to Flash.
For capacity expansion beyond 2,500 concurrent connections, submit a request to DeepSeek at no additional cost.
Routing Configuration Update
If you use TheRouter to route requests, update your provider configuration to use the canonical model name.
# Before
providers:
- id: deepseek
models:
- deepseek-v4-pro
- deepseek-v4-flash
# After
providers:
- id: deepseek
models:
- deepseek-flash # V4.1 Flash — replaces both V4 Pro and V4 Flash
If you had separate routing rules for V4 Pro (quality tier) and V4 Flash (cost tier), you now have a single model. Consider adding a fallback chain that routes to another provider's Pro-tier model when you need guaranteed Pro-level quality.
- Swap three values, not three SDKs. Change
api_key,base_url, andmodelin the existing OpenAI client. Keep your request/response code unchanged. - Map model IDs explicitly. The target provider's model id is almost never identical to the OpenAI id. Keep a single dict of
{ openai_id: target_id }outside business logic. - Verify streaming format. SSE chunks must follow the OpenAI
data: {...}+data: [DONE]contract. Test one streaming call before moving production traffic. - Check rate-limit headers. Some providers omit
x-ratelimit-*headers. Add a wrapper that defaults safely when headers are absent. - Keep a rollback path. Ship the swap behind a feature flag, run both endpoints in shadow for 24 hours, then cut over.
Rollback Plan
There is no rollback to V4 Pro after September 14. DeepSeek will reroute all V4 Pro requests to V4.1 Flash unconditionally.
If V4.1 Flash does not meet your quality requirements for Pro-tier workloads, your options are:
- Wait for V4.1 Pro. DeepSeek has confirmed a V4.1 Pro release but has not announced a date or pricing. Until then, the Pro tier is unavailable on DeepSeek.
- Route to another provider. For workloads that need frontier-class reasoning, consider routing to Claude Opus 4, GPT-6 Astra, or Kimi K3 as a primary or fallback target.
- Test V4.1 Flash now. Run your eval suite against
deepseek-flashbefore September 14. If it passes, you save 70–86% on cost with higher concurrency. That is the likely outcome for most workloads.
Existing TheRouter Blog Posts Affected
Five of our published blog posts reference deepseek-v4-pro by name. After September 14, the model name in those posts will resolve to V4.1 Flash. We will update those posts to reflect the new model landscape:
- DeepSeek API: The Complete Guide
- DeepSeek V4 Pro GA Peak Pricing Routing Guide
- DeepSeek V4 Pro vs Flash API Comparison
- Kimi K3 vs DeepSeek V4 vs Qwen3.7 Comparison
- Qwen3.8-Max vs DeepSeek V4 Pro Comparison
FAQ
Will deepseek-v4-pro stop working on September 14?
No. Requests will continue to be accepted. They will be served by V4.1 Flash at Flash pricing. Your API calls will not break, but the underlying model changes.
Is V4.1 Flash the same as the old V4 Flash? No. V4.1 Flash is a new model (DeepSeek-V4.1-Flash). The old V4 Flash (DeepSeek-V4-Flash) has already been retired and its name redirects to V4.1 Flash.
When will V4.1 Pro be available? DeepSeek has confirmed a V4.1 Pro release but has not announced a date. The pricing page states that V4 Pro will route to V4.1 Flash "until V4.1 Pro is released in the future."
Do I need to change my API key? No. The same API key works. Only the model name needs updating.
Does this affect DeepSeek R1? No. DeepSeek R1 is a separate model line and is not part of this consolidation.