DashScope October 2026 Model Sunset: 27-Day Last-Call Migration Checklist
October 10 is 27 days away. DashScope (Bailian) will hard-stop 100+ models — Qwen mainlines, snapshots, and every third-party model (DeepSeek, GLM, MiniMax, Kimi). QPM/TPM throttling is already underway. Here is the compact last-call checklist: audit your model IDs, swap replacements, get vendor-direct keys, and test before the deadline.
DashScope October 2026 Model Sunset: 27-Day Last-Call Migration Checklist
On October 10, 2026 at 00:00 CST, DashScope (Bailian) will hard-stop every model in its largest-ever retirement wave. No grace period. No degraded mode. QPM and TPM throttling has already started for retiring models. If you are still calling any of these model IDs, the clock is running.
We published a comprehensive migration guide in July with the full model list, pricing analysis, and before/after code diffs. This post is not that guide. This is the compact, no-extras checklist for the final 27 days.
- Swap three values, not three SDKs. Change
api_key,base_url, andmodelin the existing OpenAI client. Keep your request/response code unchanged. - Map model IDs explicitly. The target provider's model id is almost never identical to the OpenAI id. Keep a single dict of
{ openai_id: target_id }outside business logic. - Verify streaming format. SSE chunks must follow the OpenAI
data: {...}+data: [DONE]contract. Test one streaming call before moving production traffic. - Check rate-limit headers. Some providers omit
x-ratelimit-*headers. Add a wrapper that defaults safely when headers are absent. - Keep a rollback path. Ship the swap behind a feature flag, run both endpoints in shadow for 24 hours, then cut over.
What Is Retiring — The 30-Second Summary
Six official notices, one deadline. Everything retires October 10:
| Notice | What Dies | Count |
|---|---|---|
| #118344 | Qwen3 mainlines: qwen3-max, qwen3-max-preview, qwen3.6-max-preview, qwen3-coder-plus, qwen3-vl-flash | 5 |
| #118177 | Legacy mainlines: qwen-turbo, qwen-vl-max/plus, qwq-plus, qvq-max/plus, qwen-math-turbo, qwen-coder-turbo/plus | 10 |
| #118434 | Legacy models: older Qwen mainlines | varies |
| #118345 | 60+ snapshots + ALL third-party: DeepSeek V3/R1, MiniMax-M2.1, GLM-4.7/4.6, Kimi-K2 | 60+ |
| #118331 | Voice mainlines (cosyvoice, paraformer, gummy) | 15+ |
| #118332 | Voice snapshots | varies |
The third-party model removal is the hardest migration. DashScope suggests qwen3.7-plus as a generic replacement for everything, but if you picked DeepSeek R1 for reasoning or GLM-4.7 for Chinese long-context, swapping to a different model family is not a migration. It is a rewrite. Switch to vendor-direct instead.
The Checklist — 5 Steps, Do Them Now
Step 1: Audit Your Model Usage (Day 1)
Go to the DashScope model telemetry page and check if any of your active endpoints call a retiring model ID.
Then grep your codebase:
grep -rn "qwen3-max\|qwen3-max-preview\|qwen3.6-max-preview\|qwen3-coder-plus\|qwen3-vl-flash\|qwen-turbo\|qwen-vl-max\|qwen-vl-plus\|qwq-plus\|qvq-max\|qvq-plus\|qwen-math-turbo\|qwen-coder-turbo\|qwen-coder-plus\|deepseek-v3\|deepseek-r1\|MiniMax-M2.1\|glm-4.7\|glm-4.6\|Moonshot-Kimi-K2\|kimi-k2-thinking" \
--include="*.py" --include="*.ts" --include="*.js" --include="*.yaml" --include="*.yml" --include="*.json" .
If nothing comes back, you are clear. Stop reading.
Step 2: Swap Qwen Model IDs (Days 2–5)
For Qwen-family models, the migration is a model ID swap. Same endpoint, same API key, same SDK.
Quick-reference replacement map:
| Retiring | Replacement | Price Change |
|---|---|---|
qwen3-max | qwen3.7-max | +380% input at <32K, -41% at 128K+ |
qwen3-max-preview | qwen3.7-max | same as above |
qwen3.6-max-preview | qwen3.7-max | see pricing page |
qwen3-coder-plus | qwen3.7-plus | marginal |
qwen3-vl-flash | qwen3.6-flash | marginal |
qwen-turbo | qwen3.6-flash | -67% (cheaper) |
qwq-plus | qwen3.7-plus + enable_thinking: true | API surface change |
qvq-max / qvq-plus | qwen3.7-plus + enable_thinking: true | API surface change |
qwen-coder-turbo / qwen-coder-plus | qwen3.7-plus | marginal |
If you used reasoning models (qwq-plus, qvq-max, qvq-plus), the replacement requires an API surface change. The Qwen3.7 series uses enable_thinking: true and returns reasoning in a reasoning_content field. Thinking tokens are billed at the output rate. Read our Qwen3.7 series guide for the thinking-mode API.
Budget-sensitive teams: If qwen3.7-max pricing (¥12/¥36 per million tokens) is too steep for your workload, try qwen3.8-max at the same price point but with a newer model, or step down to qwen3.7-plus at ¥2/¥8. Also consider the new qwen3.8-max-prime fast mode at ¥24/¥72 if latency matters more than cost.
Step 3: Get Vendor-Direct API Keys for Third-Party Models (Days 2–5)
Every third-party model hosted on DashScope dies on October 10. There is no Qwen-based replacement that makes sense for these. Switch to vendor-direct:
| Retiring DashScope Model | Vendor | Vendor-Direct Endpoint | Guide |
|---|---|---|---|
deepseek-v3, deepseek-v3.1, deepseek-v3.2 | DeepSeek | https://api.deepseek.com | DeepSeek API guide |
deepseek-r1, deepseek-r1-0528, distill variants | DeepSeek | https://api.deepseek.com | DeepSeek API guide |
glm-4.7, glm-4.6 | Zhipu AI | https://open.bigmodel.cn | Zhipu GLM API guide |
Moonshot-Kimi-K2-Instruct, kimi-k2-thinking | Moonshot AI | https://api.moonshot.cn | Kimi K3 API guide |
MiniMax-M2.1 | MiniMax | https://api.minimax.chat | — |
For open-source models (DeepSeek V3, R1, R1-distill variants), you can also route through SiliconFlow which hosts them at competitive rates. See our SiliconFlow free models guide for details.
The code change for vendor-direct:
# Before — DeepSeek through DashScope (dying October 10)
client = OpenAI(
api_key="sk-dashscope-xxx",
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1"
)
response = client.chat.completions.create(
model="deepseek-v3",
messages=[{"role": "user", "content": "Hello"}]
)
# After — DeepSeek vendor-direct
client = OpenAI(
api_key="sk-deepseek-xxx",
base_url="https://api.deepseek.com"
)
response = client.chat.completions.create(
model="deepseek-chat", # vendor's own model ID
messages=[{"role": "user", "content": "Hello"}]
)
Step 4: Test Replacements (Days 5–14)
Run your evaluation suite against the replacement model. Focus on three things:
- Output format. Reasoning models (
qwq-plustoqwen3.7-pluswith thinking) return reasoning in a different field. Structured output formats may differ between model families. - Context window.
qwen3-maxhad tiered pricing across 256K.qwen3.7-maxhas flat-rate 1M context. Your prompts will work, but your bill changes. - Latency and rate limits. DashScope is actively throttling retiring models. Your baseline latency measurements are unreliable if you run them on the old model IDs.
Step 5: Deploy and Monitor (Days 14–27)
Push updated model IDs and base URLs to production. Then:
- Watch your error rates. If DashScope further reduces QPM/TPM on retiring models during the final weeks, your old endpoints may start returning 429s before October 10.
- Set up alerts on model ID strings. If any service still references a retiring model after your deploy, you want to catch it before the hard stop.
- If you route through TheRouter, configure model fallback from retiring IDs to replacements. TheRouter routes OpenAI-compatible requests through configured providers and supports provider/model routing and fallback where the live product path supports it.
QPM/TPM Throttling — Already In Progress
DashScope started reducing rate limits on retiring models from the day each notice was published. The earliest notice (#118177) was published April 13, 2026 — QPM and TPM for those models have been declining for five months. If you had expanded limits, they were reset to defaults first, then reduced below defaults (Deprecation mechanism, retrieved 2026-09-13).
If you are seeing increased 429 responses or slower throughput on any of the retiring models, this is the throttling. It will only get worse between now and October 10.
What Happens on October 10 at 00:00 CST
- Model inference stops. API calls to retiring model IDs will fail. No fallback, no degraded output.
- Fine-tuned models based on retired base models may still function if already deployed, but new fine-tuning on retired bases is blocked.
- Console and documentation for retired models go offline.
- No extension mechanism. DashScope deadlines have historically been extended (the original September 8 Qwen3 deadline moved to October 10), but counting on another extension is a gamble.
Related Migration Resources
For comprehensive model-by-model mapping with pricing analysis and before/after code diffs, see our full October 2026 DashScope migration guide.
For the Qwen3 to Qwen3.7 upgrade path specifically, see our Qwen3 to Qwen3.7 upgrade guide.
For the DashScope model lifecycle reference with all past and upcoming waves, see our DashScope model lifecycle consolidation.
Sources: Aliyun Model Deprecation Page (retrieved 2026-09-13), Model Pricing (retrieved 2026-09-13), Notice #118344, Notice #118177, Notice #118345, Notice #118434, Notice #118331, Notice #118332.