Back to Models

Qwen Max

qwenqwen/qwen-max

Previous-gen Qwen flagship β€” still strong on reasoning and tool use.

qwen-max is the previous-generation flagship in the Qwen series. Even after Qwen3 launched it remains a credible choice for reasoning-heavy chat, structured output, and function-calling workloads, with a 128K context window that comfortably handles long documents.

Use it when you want flagship-tier behaviour at a flat, predictable price and don't need Qwen3-level coding or the 256K context of qwen3-max. TheRouter routes between bailian-cn and bailian-sg per request based on cost.

Strong reasoning
Solid chain-of-thought, math, and multi-step tool-use performance from the previous-gen Qwen flagship.
128K context
Long-document analysis, large code reviews, and extended agent histories fit in a single call.
Function calling
Native tool/function calling with OpenAI-compatible schemas β€” no wrapper SDK needed.
Dual-region routing
Selector picks bailian-cn or bailian-sg per request based on cost β€” failover is automatic.
When to use
Reasoning, structured output, and tool-use workloads where you want a flagship-tier Qwen at a steady price and Qwen3-max's premium is not justified.
When not to use
If you need the absolute top of Qwen on coding or 256K context, upgrade to qwen3-max. For high-volume cheap chat, drop to qwen-plus / qwen-turbo.
Pricing: $2.00 input / $8.00 output per MTok. TheRouter picks the cheaper of bailian-cn / bailian-sg per request.
Context Length
131K
Max Output
33K
Input Priceper 1M tokens
$1.73/ 1M tokens
Output Priceper 1M tokens
$6.91/ 1M tokens

Modalities

text→text

Pricing Breakdown

TypeRate
Input$1.73 / 1M tokens
Output$6.91 / 1M tokens

Supported Parameters

temperaturemax_tokenstop_ptoolstool_choiceresponse_formatstop

Specifications

ArchitectureLarge-scale Mixture-of-Experts (MoE) β€” exact active/total parameter counts not publicly disclosedqwenlm.github.io β†—to verify
Pretraining dataOver 20 trillion tokensqwenlm.github.io β†—verified
Post-trainingSupervised Fine-Tuning (SFT) + Reinforcement Learning from Human Feedback (RLHF)qwenlm.github.io β†—verified
API availability date (Qwen2.5-Max)January 28, 2025 (pinned alias: qwen-max-2025-01-25)qwenlm.github.io β†—verified
Thinking modeNot supported β€” this generation predates Qwen3's hybrid thinking API. Use qwen/qwen3-32b or newer for thinking mode.verified
LicenseProprietary (closed-weight, API-only)verified
Lifecycle statusLegacy β€” Alibaba classifies qwen-max as a legacy endpoint; new projects should use Qwen3.x modelshelp.aliyun.com β†—verified

Benchmarks

BenchmarkDistributionScoreSource
Arena-Hard
Qwen2.5-Max outperformed DeepSeek V3, GPT-4o, and Claude-3.5-Sonnet on Arena-Hard at launch according to the official Qwen2.5-Max blog. Exact numeric score not reported.
SOTA at launch (Jan 2025)rankqwenlm.github.io β†—
LiveCodeBench
Qwen2.5-Max outperformed DeepSeek V3 on LiveCodeBench at launch. Specific pass@1 score not reported in the official blog.
Above DeepSeek V3 at launchrankqwenlm.github.io β†—
GPQA-Diamond
Qwen2.5-Max outperformed DeepSeek V3 on GPQA-Diamond at launch. Exact score not published in the official blog.
Above DeepSeek V3 at launchrankqwenlm.github.io β†—

API Usage Examples

Use the global api.therouter.ai endpoint shown below for new integrations; the legacy China accelerated endpoint is retired.

cURL
curl https://api.therouter.ai/v1/chat/completions   -H "Content-Type: application/json"   -H "Authorization: Bearer $THE_ROUTER_API_KEY"   -d '{
    "model": "qwen/qwen-max",
    "messages": [
      {"role": "user", "content": "Summarize the key points from this input."}
    ]
  }'

Chat completion

Standard OpenAI-compatible chat completions. If you are building something new, consider qwen/qwen3-32b instead β€” same provider family, active generation, with optional thinking mode.

cURL
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen/qwen-max",
    "messages": [
      {"role": "user", "content": "Summarize the key capabilities of Qwen2.5-Max."}
    ],
    "temperature": 0.7,
    "max_tokens": 1024
  }'

More from qwen

Similar models

Cross-provider sibling models

News & changes

2025-01-28

Qwen2.5-Max launches as qwen-max API alias, outperforming DeepSeek V3 at launch

Alibaba released Qwen2.5-Max via the qwen-max DashScope alias, reporting benchmark wins over DeepSeek V3 on Arena-Hard, LiveBench, LiveCodeBench, and GPQA-Diamond. The model is a large-scale MoE pretrained on over 20 trillion tokens with SFT and RLHF post-training.

re-authored by TheRouterQwen blog β†—

Frequently asked

What is qwen/qwen-max on TheRouter?

qwen-max is a floating commercial alias on Alibaba DashScope that has pointed to Qwen2.5-Max β€” a large-scale MoE model β€” since January 2025. Alibaba now classifies it as a legacy endpoint; the current recommended models are in the Qwen3.x family.

Should I use qwen/qwen-max for a new project?

No. Use qwen/qwen3-32b for balanced general tasks, qwen/qwen3-coder-480b for coding agents, or qwen/qwen3-235b for highest-quality reasoning. qwen-max is best kept for teams with existing integrations that are not yet ready to migrate.

Does qwen/qwen-max support thinking mode?

No. qwen-max predates Qwen3's hybrid thinking API and does not support the reasoning parameter. For thinking-mode reasoning, use qwen/qwen3-32b or qwen/qwen3-235b.

How does qwen/qwen-max pricing compare to Qwen3 models?

qwen/qwen-max is priced at $2.00/MTok input and $8.00/MTok output on TheRouter. qwen/qwen3-32b is priced lower and offers hybrid thinking mode and a 131K context window β€” making it the better value for most new workloads.

Fact ledger β€” every claim on this page traces here
sourceURLretrieved
Architectureqwenlm.github.io β†—2026-06-10to verify
Pretraining dataqwenlm.github.io β†—2026-06-10verified
Post-trainingqwenlm.github.io β†—2026-06-10verified
API availability date (Qwen2.5-Max)qwenlm.github.io β†—2026-06-10verified
Thinking modeβ€”β€”verified
Licenseβ€”β€”verified
Lifecycle statushelp.aliyun.com β†—2026-06-10verified
Arena-Hardqwenlm.github.io β†—2026-06-10to verify
LiveCodeBenchqwenlm.github.io β†—2026-06-10to verify
GPQA-Diamondqwenlm.github.io β†—2026-06-10to verify
Qwen2.5-Max launches as qwen-max API alias, outperforming DeepSeek V3 at launchQwen blog β†—2026-06-10verified
Help & contact