Back to Models

qwen3-max is Alibaba's flagship commercial model in the Qwen3 generation. It leads the Qwen lineup on reasoning, coding, math, and complex tool use, with a 256K context window for repo-scale code, long documents, and multi-turn agent traces.

TheRouter ships this model on both the mainland (bailian-cn) and Singapore (bailian-sg) Bailian endpoints. The selector picks the cheaper region per request so you get vendor-direct latency without manual region pinning.

Deep reasoning
Best-in-family reasoning, math, and tool-use scores. Comfortable with chain-of-thought-heavy prompts.
256K context
Fits a small repo or several long PDFs in a single request β€” no chunking gymnastics required.
Dual-region routing
Auto-selects between bailian-cn and bailian-sg per request based on cost β€” no client-side region logic.
OpenAI-compatible
Standard `/v1/chat/completions` shape including tools, JSON mode, and streaming.
When to use
High-stakes reasoning, complex coding tasks, repo-aware refactors, and agent workloads where Qwen-family quality matters and budget allows for a flagship-tier model.
When not to use
High-volume customer support, simple rewrites, or latency-critical autocomplete β€” drop down to qwen-plus or qwen-turbo for 10–20Γ— lower cost.
Pricing: $1.50 input / $7.50 output per MTok. TheRouter routes to the cheaper of bailian-cn / bailian-sg per request.
Context Length
262K
Max Output
33K
Input Priceper 1M tokens
$1.30/ 1M tokens
Output Priceper 1M tokens
$6.48/ 1M tokens

Modalities

text→text

Pricing Breakdown

TypeRate
Input$1.30 / 1M tokens
Output$6.48 / 1M tokens

Supported Parameters

temperaturemax_tokenstop_ptoolstool_choiceresponse_formatstop

Specifications

Stable aliasqwen3-max β†’ qwen3-max-2026-01-23www.alibabacloud.com β†—verified
Available snapshotsqwen3-max-2026-01-23; qwen3-max-preview; qwen3-max-2025-09-23www.alibabacloud.com β†—verified
ArchitectureQwen3 Mixture-of-Experts (MoE) with over 1 trillion total parameterswww.alibabacloud.com β†—verified
Pretraining data36 trillion tokenswww.alibabacloud.com β†—verified
Context window256K tokens upstreamwww.alibabacloud.com β†—verified
Max output32,768 tokens in TheRouter configgithub.com β†—verified
Thinking modeHybrid thinking is supported upstream for qwen3-max and qwen3-max-2026-01-23; thinking is disabled by default. TheRouter's qwen/qwen3-max public config does not currently expose a reasoning parameter.www.alibabacloud.com β†—verified
API capabilitiesOpenAI-compatible chat, streaming, function calling, structured output / JSON mode, and stop sequences through TheRouter; built-in Alibaba tools are documented upstream but not represented as separate TheRouter endpoints.www.alibabacloud.com β†—verified
TheRouter price bands$1.20/M input + $6.00/M output up to 32K; $2.40/M input + $12.00/M output after 32K. Alibaba also publishes a third 128K-256K band ($3/$15) that TheRouter's current single-threshold schema cannot fully encode.www.alibabacloud.com β†—verified
LicenseProprietary (closed-weight, API-only)verified

Benchmarks

BenchmarkDistributionScoreSource
SWE-bench Verified
Alibaba reports Qwen3-Max-Instruct at 69.6 on SWE-bench Verified, positioning it among top coding models at launch.
69.6scorealibabacloud.com β†—
Tau2-Bench
Alibaba describes Tau2-Bench as an evaluation of agent tool-calling proficiency and says Qwen3-Max-Instruct surpassed Claude Opus 4 and DeepSeek V3.1 on this benchmark.
74.8scorealibabacloud.com β†—
LMArena Text Arena
The launch article says Qwen3-Max-Instruct preview ranked third on the Text Arena leaderboard and surpassed GPT-5-Chat. A direct leaderboard snapshot is not embedded in the article, so this is treated as an official claim rather than a reproduced leaderboard row.
Top-three global ranking for previewrankalibabacloud.com β†—
AIME 25 / HMMT
This is explicitly for the still-training Qwen3-Max-Thinking variant augmented with a code interpreter and parallel test-time compute, not for the default non-thinking qwen3-max call path.
100% with Qwen3-Max-Thinking + tools + scaled test-time compute%alibabacloud.com β†—

API Usage Examples

Use the global api.therouter.ai endpoint shown below for new integrations; the legacy China accelerated endpoint is retired.

cURL
curl https://api.therouter.ai/v1/chat/completions   -H "Content-Type: application/json"   -H "Authorization: Bearer $THE_ROUTER_API_KEY"   -d '{
    "model": "qwen/qwen3-max",
    "messages": [
      {"role": "user", "content": "Summarize the key points from this input."}
    ]
  }'

Chat completion

Use Qwen3-Max through TheRouter's OpenAI-compatible Chat Completions endpoint for difficult text, coding, and agent prompts.

cURL
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen/qwen3-max",
    "messages": [{"role": "user", "content": "Summarize this vendor contract and list risky clauses."}],
    "temperature": 0.2,
    "top_p": 0.8,
    "max_tokens": 4096
  }'

More from qwen

Similar models

Cross-provider sibling models

News & changes

2026-01-23

qwen3-max stable alias advances to the January 2026 snapshot

Alibaba Cloud's pricing table identifies qwen3-max as currently equivalent to qwen3-max-2026-01-23, with non-thinking and thinking modes available for the current snapshot and tiered billing up to 256K tokens.

re-authored by TheRouteralibabacloud.com β†—
2025-09-23

Qwen3-Max launches as Alibaba's trillion-parameter Qwen flagship

The Qwen team introduced Qwen3-Max with over 1 trillion parameters, 36 trillion pretraining tokens, OpenAI-compatible API access as qwen3-max, and reported gains on coding and agent benchmarks.

re-authored by TheRouteralibabacloud.com β†—

Frequently asked

Is qwen/qwen3-max the same as Qwen3-Max-Instruct?

For TheRouter users, qwen/qwen3-max is the public id that routes to Alibaba's qwen3-max chat/instruct API. Alibaba's launch article names the API model qwen3-max for Qwen3-Max-Instruct, while the pricing page says the stable alias now maps to qwen3-max-2026-01-23.

re-authored by TheRouterwww.alibabacloud.com β†—
Does qwen/qwen3-max support thinking mode on TheRouter?

Upstream Alibaba documents hybrid thinking for qwen3-max and qwen3-max-2026-01-23, disabled by default. TheRouter's current public model config for qwen/qwen3-max does not expose reasoning, so treat TheRouter calls as the standard non-thinking chat path unless your account has a provider-specific override.

re-authored by TheRouterwww.alibabacloud.com β†—
How much context does qwen/qwen3-max support?

The upstream Alibaba model list marks qwen3-max as 256K context, and TheRouter config sets qwen/qwen3-max to 262,144 tokens with 32,768 max completion tokens. If you need 1M context, compare qwen/qwen3.7-max or qwen/qwen3.7-plus instead.

re-authored by TheRouterwww.alibabacloud.com β†—
Should I choose qwen3-max or qwen3.7-max?

For new maximum-capability deployments, start with qwen3.7-max because Alibaba's current guide recommends it for the strongest reasoning and 1M context. Keep qwen3-max when you need the 256K class, existing snapshot behavior, or current TheRouter pricing/routing stability.

re-authored by TheRouterwww.alibabacloud.com β†—
Fact ledger β€” every claim on this page traces here
sourceURLretrieved
Stable aliaswww.alibabacloud.com β†—2026-07-30verified
Available snapshotswww.alibabacloud.com β†—2026-07-30verified
Architecturewww.alibabacloud.com β†—2026-07-30verified
Pretraining datawww.alibabacloud.com β†—2026-07-30verified
Context windowwww.alibabacloud.com β†—2026-07-30verified
Max outputgithub.com β†—2026-07-30verified
Thinking modewww.alibabacloud.com β†—2026-07-30verified
API capabilitieswww.alibabacloud.com β†—2026-07-30verified
TheRouter price bandswww.alibabacloud.com β†—2026-07-30verified
Licenseβ€”β€”verified
SWE-bench Verifiedalibabacloud.com β†—2026-07-30verified
Tau2-Benchalibabacloud.com β†—2026-07-30verified
LMArena Text Arenaalibabacloud.com β†—2026-07-30to verify
AIME 25 / HMMTalibabacloud.com β†—2026-07-30to verify
qwen3-max stable alias advances to the January 2026 snapshotalibabacloud.com β†—2026-07-30verified
Qwen3-Max launches as Alibaba's trillion-parameter Qwen flagshipalibabacloud.com β†—2026-07-30verified
Is qwen/qwen3-max the same as Qwen3-Max-Instruct?www.alibabacloud.com β†—2026-07-30to verify
Does qwen/qwen3-max support thinking mode on TheRouter?www.alibabacloud.com β†—2026-07-30to verify
How much context does qwen/qwen3-max support?www.alibabacloud.com β†—2026-07-30to verify
Should I choose qwen3-max or qwen3.7-max?www.alibabacloud.com β†—2026-07-30to verify
Help & contact