Aliyun Bailian API: Qwen, Wan, DashScope Dual-Region Routing

Looking for the Aliyun Bailian API, the DashScope OpenAI-compatible base_url, or Qwen / text-embedding-v4 pricing? TheRouter connects one API key to Alibaba Cloud Model Studio in Beijing (dashscope.aliyuncs.com) and Singapore (dashscope-intl.aliyuncs.com), exposes them behind https://api.therouter.ai/v1, and auto-routes each request to the cheaper region for that model.


Why this matters for 阿里云百炼 API users

The same Qwen model is priced differently on Alibaba’s domestic console and its international console — and which one wins varies model by model. Routing through an aggregator forfeits that arbitrage, adds a markup, and hides the provider’s real rate-limit responses. TheRouter ships a single provider-aliyun-bailian service deployed twice (Beijing + Singapore), each with its own DashScope API key, and the router compares the per-MTok cost of every active route on every request. For Qwen today the China region is cheaper across the board, so it ranks first; if that ever inverts on a single model, Singapore takes over automatically with no config change.

We also keep enable_thinking, reasoning_content, and the rest of Qwen’s OpenAI-incompatible extensions passing through verbatim — no rewriting, no swallowed fields.

Model coverage at launch

FamilyModelsUse case
Qwen3 commercialqwen3-max · qwen-max · qwen-plus · qwen-flashFlagship to ultra-cheap general chat
Long context (CN-only)qwen-long10M-token document analysis
Codeqwen3-coder-plusRepo-aware coding, function calling
Visionqwen3-vl-plusText + image input, 256K context
Reasoning (thinking)qwen-plus-thinkingFull thinking traces preserved
Embeddingstext-embedding-v3 · text-embedding-v4Multilingual, 8K context
Wan 2.2 (image)wan2.2-t2i-flash · wan2.2-t2i-plusText → image, async via /v1/jobs

Full 13-model list with pricing on the Aliyun Bailian provider page. Wan text-to-video (wan2.2-t2v) is intentionally not routed at launch — DashScope bills it per second, and our gateway only supports flat per-request billing. It will come back once duration-aware billing ships.

Pricing

All prices are per million tokens unless noted, and routed to whichever region is cheaper for that specific model. If you call Aliyun Bailian directly, the usual OpenAI-compatible endpoints aredashscope.aliyuncs.com/compatible-mode/v1for Beijing anddashscope-intl.aliyuncs.com/compatible-mode/v1for Singapore; through TheRouter, both sit behindhttps://api.therouter.ai/v1.

  • qwen/qwen-flash — $0.07 in / $0.50 out (the cheapest serious chat we route today)
  • qwen/qwen-plus — $0.50 in / $1.50 out (mid-tier with thinking mode, 1M context)
  • qwen/qwen3-max — $1.50 in / $7.50 out (current flagship)
  • qwen/qwen-plus-thinking — $0.26 in / $2.69 out (reasoning, thinking traces preserved)
  • qwen/qwen3-vl-plus — $0.30 in / $2.00 out (vision)
  • qwen/text-embedding-v4 — $0.12 in
  • wan/wan2.2-t2i-flash — $0.04/image · wan/wan2.2-t2i-plus — $0.08/image

Prices shown are launch-time figures; see each model page for the current billed rate.

Three things to know about Bailian

  1. Region selection is automatic, per-model. You never pick “CN” or “SG” — you ask for qwen/qwen-plus and the router picks the cheaper region for that specific model on that specific request. If China prices a model lower this week and Singapore prices it lower next week, your bill follows without a code change.
  2. Thinking mode passes through. qwen-plus and qwen-plus-thinking accept enable_thinking and return reasoning_content in the response — no translation, no munging. If you’ve used Qwen’s native API, the response shape will feel identical.
  3. Wan image gen is async. Wan goes through TheRouter’s /v1/jobs surface — submit a job, poll until completed, fetch the image URL. Same pattern as GPT-Image-2, SeedDance, and CogVideoX.

Getting started

Plain chat — swap the model ID, that’s it:

curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THE_ROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen/qwen-flash",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

Thinking mode with QwQ:

curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THE_ROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen/qwen-plus-thinking",
    "messages": [{"role": "user", "content": "Why is the sky blue?"}],
    "extra_body": {"enable_thinking": true}
  }'
# response.choices[0].message.reasoning_content holds the thinking trace

Vision (image + text → text):

curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THE_ROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen/qwen3-vl-plus",
    "messages": [{"role": "user", "content": [
      {"type": "text", "text": "Describe this image"},
      {"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}}
    ]}]
  }'

Multilingual embeddings:

curl https://api.therouter.ai/v1/embeddings \
  -H "Authorization: Bearer $THE_ROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen/text-embedding-v4",
    "input": ["The quick brown fox jumps over the lazy dog"]
  }'

Wan text-to-image (async):

# 1. Submit
curl https://api.therouter.ai/v1/jobs \
  -H "Authorization: Bearer $THE_ROUTER_API_KEY" \
  -d '{
    "operation": "images.generate",
    "model": "wan/wan2.2-t2i-flash",
    "prompt": "A neon-lit Shanghai skyline at dusk, photorealistic"
  }'
# → {"id": "img_xxx", "status": "in_progress"}

# 2. Poll until completed
curl https://api.therouter.ai/v1/jobs/img_xxx \
  -H "Authorization: Bearer $THE_ROUTER_API_KEY"
# → {"id": "...", "status": "completed", "image_url": "https://..."}

Featured model detail pages

Help & contact