Qwen API via OpenAI SDK: Complete Integration Guide for DashScope (2026)
How to call Qwen models through DashScope's OpenAI-compatible endpoint — base URL, model IDs, thinking mode, tool calling, pricing, and routing through TheRouter.
DashScope serves the entire Qwen model family through an endpoint that speaks the OpenAI Chat Completions protocol. You change three values in your existing OpenAI SDK code — API key, base URL, model name — and the same application that talked to GPT-4o now talks to Qwen3.8-Max, Qwen3.7-Plus, or any of the third-party models hosted on DashScope (DeepSeek V4, Kimi K3, GLM-5.3). No request-format rewriting, no new client library.
This guide covers every step from "I have no Alibaba Cloud account" to "I am routing Qwen alongside Claude and GPT through a single gateway in production."
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
Getting Started in 3 Minutes
Step 1 — Create an Alibaba Cloud account. Go to alibabacloud.com for international access or aliyun.com for China mainland. Activate Model Studio (百炼) from the console.
Step 2 — Generate an API key. Open the API Keys page inside Model Studio and create a new key. Keys are region-scoped: a key created in the Beijing region only works against Beijing endpoints, and a Singapore key only works against Singapore endpoints.
Step 3 — Install the OpenAI SDK.
pip install --upgrade openai
Step 4 — Make your first call.
from openai import OpenAI
import os
client = OpenAI(
api_key=os.getenv("DASHSCOPE_API_KEY"),
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
)
response = client.chat.completions.create(
model="qwen3.8-max",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain API routing in two sentences."},
],
)
print(response.choices[0].message.content)
The response object is identical to what openai.ChatCompletion returns — same id, choices, usage fields. Any code that parses OpenAI responses works without modification.
Base URL by Region
DashScope operates in multiple regions, each with its own endpoint. Alibaba Cloud has been migrating to workspace-specific domain names that provide better performance and higher stability.
| Region | Base URL |
|---|---|
| Beijing (China) | https://dashscope.aliyuncs.com/compatible-mode/v1 |
| Beijing (workspace) | https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/compatible-mode/v1 |
| Singapore | https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1 |
| Virginia (US) | https://dashscope-us.aliyuncs.com/compatible-mode/v1 |
| Hong Kong | https://{WorkspaceId}.cn-hongkong.maas.aliyuncs.com/compatible-mode/v1 |
| Tokyo | https://{WorkspaceId}.ap-northeast-1.maas.aliyuncs.com/compatible-mode/v1 |
Replace {WorkspaceId} with the ID shown in your Model Studio console. The legacy dashscope.aliyuncs.com endpoint still works for Beijing but Alibaba Cloud recommends the workspace-specific domain for all new integrations.
A common first-time error is using a Beijing API key against the Singapore endpoint (or vice versa). DashScope returns HTTP 401 invalid_api_key when the key and endpoint regions do not match.
Model IDs: The Current Qwen Lineup
DashScope organizes models by generation and tier. Here are the text-generation models most relevant to API developers as of August 2026.
Qwen Max Tier (Flagship)
| Model ID | Parameters | Context | Key Capabilities |
|---|---|---|---|
qwen3.8-max | 2.4T MoE | 1M | Flagship. Native vision. Extended coding agent sessions. |
qwen3.8-max-prime | 2.4T MoE | 1M | Speed-optimized variant of qwen3.8-max. |
qwen3.7-max | Undisclosed | 1M | Previous flagship. Thinking + non-thinking modes. Currently at 50% promotional pricing. |
Qwen Plus Tier (Balanced)
| Model ID | Context | Notes |
|---|---|---|
qwen3.7-plus | 1M | Strong multimodal with text, image, video input at lower cost. |
qwen-plus | 128K | Legacy Plus tier. Still available. |
Qwen Flash / Turbo Tier (Speed + Cost)
| Model ID | Context | Notes |
|---|---|---|
qwen3.8-flash | 1M | Latest Flash model. Fast inference, multimodal. |
qwen3.7-flash | 1M | Multimodal Flash with vision. |
qwen-turbo | 1M | Budget option for high-throughput workloads. |
Open-Source Models on DashScope
| Model ID | Parameters | Notes |
|---|---|---|
qwen3.8-2.4t-a95b | 2.4T/95B active | Open-source MoE flagship (August 2026). |
qwen3.8-27b | 27B dense | Compact model with good coding performance. |
qwen3-8b | 8B | Lightweight, suitable for embedding in applications. |
Third-Party Models on DashScope
DashScope is not Qwen-only. Alibaba Cloud hosts models from other providers under the same OpenAI-compatible endpoint, so you can switch between vendors by changing only the model ID.
| Provider | Model ID on DashScope | Category |
|---|---|---|
| Moonshot AI | kimi-k3 | Text generation, reasoning (2.8T parameters) |
| Zhipu AI | ZHIPU/GLM-5.3-Flash | Text generation, vision (320B total, 18B active) |
| Zhipu AI | ZHIPU/GLM-5.3 | Text generation, deep thinking |
| DeepSeek | deepseek-v4-pro-0813 | Text generation, reasoning (1.6T MoE) |
| DeepSeek | deepseek-v4-flash-0731 | Fast inference (284B, 13B active) |
This makes DashScope a multi-model marketplace. You get a single billing account, a single API key, and access to models from several Chinese AI labs.
Thinking Mode (Chain-of-Thought Reasoning)
Qwen3.7-Max and Qwen3.8-Max support a thinking mode where the model generates an internal reasoning trace before producing the final answer. You enable it through the extra_body parameter in the OpenAI SDK.
response = client.chat.completions.create(
model="qwen3.8-max",
messages=[
{"role": "user", "content": "Prove that the square root of 2 is irrational."},
],
extra_body={
"enable_thinking": True,
"thinking_budget": 10000,
},
stream=True,
)
When thinking mode is active, DashScope returns the reasoning trace in a separate reasoning_content field within each streamed chunk. The billing counts both the thinking tokens and the final-answer tokens.
To disable thinking on a model that defaults to it, pass enable_thinking=False in extra_body.
Tool Calling (Function Calling)
DashScope's OpenAI-compatible endpoint supports the standard tools parameter for function calling. The request and response format follows the OpenAI specification.
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a location",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City name"}
},
"required": ["location"],
},
},
}
]
response = client.chat.completions.create(
model="qwen3.8-max",
messages=[{"role": "user", "content": "What is the weather in Shanghai?"}],
tools=tools,
tool_choice="auto",
)
The model returns a tool_calls array in the assistant message, and you feed the function results back as tool role messages — identical to the OpenAI flow. Qwen3.8-Max and Qwen3.7-Max both support parallel tool calls.
Streaming
Streaming works via the standard stream=True parameter. DashScope returns chat.completion.chunk objects identical to OpenAI's format.
stream = client.chat.completions.create(
model="qwen3.8-max",
messages=[{"role": "user", "content": "Write a haiku about API routing."}],
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
delta = chunk.choices[0].delta
if delta.content:
print(delta.content, end="", flush=True)
Pass stream_options={"include_usage": True} to get token counts in the final chunk — useful for cost tracking.
Pricing (Beijing Region, August 2026)
Prices are per million tokens in RMB. Models use tiered pricing based on the total input tokens in a single request.
| Model | Input (per 1M tokens) | Output (per 1M tokens) | Free Tier |
|---|---|---|---|
| qwen3.8-max | ¥12 | ¥36 | 1M tokens |
| qwen3.8-max-prime | ¥24 | ¥72 | None |
| qwen3.7-max | ¥6 (50% promo) | ¥18 (50% promo) | 1M tokens |
| qwen3-max | ¥2.5–7 (tiered) | ¥10–28 (tiered) | 1M tokens |
Batch API calls receive 50% off on models that support it (qwen3.8-max, qwen3-max). Context caching offers a separate discount on cached input tokens — typically 10% of the standard input price for cache hits.
International regions (Singapore, Virginia, Tokyo) carry higher per-token prices. Singapore qwen3.8-max input costs ¥14.988/M tokens versus ¥12/M in Beijing.
For the latest pricing details, check the DashScope pricing page.
Common Errors and Fixes
HTTP 401 invalid_api_key — The API key and endpoint belong to different regions. Create a key in the same region as your endpoint, or switch your base URL to match your key's region.
HTTP 404 on /v1/chat/completions — You may be using the DashScope-native endpoint format instead of the OpenAI-compatible path. Make sure your base URL ends with /compatible-mode/v1, not /api/v1.
model not found — The model ID is case-sensitive. Use qwen3.8-max, not Qwen3.8-Max. Third-party model IDs use the provider prefix: ZHIPU/GLM-5.3-Flash, not glm-5.3-flash.
Thinking tokens billed but invisible — When enable_thinking is True, reasoning tokens appear in usage.completion_tokens but the reasoning trace only shows in reasoning_content during streaming. If you are not streaming, you still pay for thinking tokens but cannot inspect the trace.
Rate limit errors — DashScope applies per-model rate limits that vary by account level. These limits are not publicly documented at per-model granularity. If you hit limits on free-tier, upgrading your Alibaba Cloud account level raises the ceiling.
Production Checklist
- Use workspace-specific base URLs. The legacy
dashscope.aliyuncs.comstill works but carries no per-workspace isolation. - Rotate API keys. DashScope keys do not expire automatically. Set a rotation schedule in your secrets manager.
- Monitor the deprecation calendar. Alibaba Cloud retires older Qwen snapshots on a rolling basis. The model deprecation page lists upcoming sunsets. We track these in our DashScope model lifecycle reference.
- Enable context caching for repeated prompts. If your system prompt exceeds 1,024 tokens and is reused across requests, implicit caching activates automatically. For explicit caching, see the DashScope cache documentation.
- Test thinking mode billing impact. Thinking tokens can 2–5x your output token count. Run a representative sample before enabling in production.
- Set
stream_options.include_usageto track actual token consumption per request.
Routing Qwen Alongside Other Providers
If your application needs to hit Qwen, Claude, and GPT from a single integration point, an API router sits between your code and the upstream providers. TheRouter routes OpenAI-compatible requests to configured providers, including DashScope, and supports fallback chains so that if DashScope is down or rate-limited, your request falls through to an alternative.
A typical setup:
- Point your application at TheRouter's endpoint instead of DashScope directly.
- Configure DashScope as a provider in TheRouter with your API key and base URL.
- Set a fallback chain: try Qwen3.8-Max on DashScope first, fall back to DeepSeek V4 on the DeepSeek API if DashScope returns a 5xx.
- TheRouter forwards the request in OpenAI-compatible format, so no code changes are needed on either side.
This is particularly useful when you are already using Qwen through DashScope for primary inference but want a safety net against regional outages or temporary rate limits.
For a deeper look at DashScope's full platform capabilities, see our Aliyun Bailian API guide. For Qwen3.7-series specifics, see the DashScope Qwen3.7 series guide.
Frequently Asked Questions
Can I use the Node.js OpenAI SDK with DashScope?
Yes. Set baseURL to the DashScope compatible-mode endpoint and pass your DashScope API key. The same patterns shown in Python work identically in the Node.js SDK.
Does DashScope support the Batch API?
Yes. Models marked with "Batch调用半价" (Batch calls at 50% off) support the /v1/batches endpoint. Submit a .jsonl file of requests and retrieve results asynchronously at half the per-token cost.
What is the difference between qwen3.8-max and qwen3.8-max-prime?
Prime is a speed-optimized variant with higher per-token pricing (¥24/¥72 vs ¥12/¥36). It targets latency-sensitive workloads where faster first-token time justifies the cost premium.
Are there free credits? New accounts receive 1 million free tokens for most Qwen models, valid for 90 days from account activation or model release (whichever is later). Third-party models on DashScope may have separate promotional quotas.
Does DashScope support the Assistants API or Responses API? DashScope implements the Chat Completions endpoint. It does not support the OpenAI Assistants API or the Responses API. For agent-style orchestration, use DashScope's own Application API or build your agent loop on top of Chat Completions with tool calling.