Back to Models

Qwen Plus

qwenqwen/qwen-plus

Mid-tier Qwen with thinking + non-thinking modes and 1M context.

How TheRouter serves this differently from the vendor

As the vendor operates it

Alibaba documents the qwen-plus family on Bailian with 1M context, thinking mode support, function calling, built-in tools, and structured output. The upstream OpenAI-compatible thinking example uses the non-standard enable_thinking flag in extra_body.

On TheRouter

TheRouter serves the model as qwen/qwen-plus on the OpenAI-compatible chat route with catalog context 1,048,576, max output 32,768, and supported parameters including tools, tool_choice, response_format, reasoning, and stop. This reference treats Alibaba's 1M and feature table as upstream authority while calling out TheRouter's parameter spelling and catalog limits.

qwen-plus is the workhorse of the Qwen3 commercial line. It offers two operating modes β€” fast `enable_thinking: false` for everyday chat and a deeper `enable_thinking: true` mode that emits `reasoning_content` for harder problems β€” at a fraction of the flagship's price.

The 1M context window makes it a strong default for long-document Q&A, multi-file code reading, and long agent conversations. TheRouter routes between bailian-cn and bailian-sg per request based on cost; `enable_thinking` and `reasoning_content` pass through verbatim in both directions.

Thinking + non-thinking
Toggle `enable_thinking` per request. When true, the model returns a `reasoning_content` field alongside the final answer.
1M context
One of the largest commercial Qwen context windows β€” ideal for whole-doc or multi-file workloads.
Cost-balanced
$0.50 input / $1.50 output per MTok. ~5Γ— cheaper than qwen-max on input, ~5.3Γ— cheaper on output.
Dual-region routing
Selector picks the cheaper of bailian-cn / bailian-sg per request β€” no client-side region logic.
When to use
Default Qwen model for production: long-document Q&A, multi-step agents, code reading, and chat where you want optional deeper reasoning without paying flagship prices.
When not to use
If you need the absolute best Qwen reasoning or coding, step up to qwen3-max / qwen3-coder-plus. For high-throughput cheap chat, qwen-turbo or qwen-flash is more cost-effective.
Pricing: $0.50 input / $1.50 output per MTok. TheRouter routes to the cheaper of bailian-cn / bailian-sg per request. Thinking mode does not change the price.
Context Length
1.0M
Max Output
33K
Input Priceper 1M tokens
$0.432/ 1M tokens
Output Priceper 1M tokens
$1.30/ 1M tokens

Modalities

text→text

Pricing Breakdown

TypeRate
Input$0.432 / 1M tokens
Output$1.30 / 1M tokens

Supported Parameters

temperaturemax_tokenstop_ptoolstool_choiceresponse_formatreasoningstop

Specifications

Context length1,048,576 tokens in TheRouter catalog; Alibaba documents qwen-plus family context as 1Mhelp.aliyun.com β†—verified
Maximum output32,768 tokens in TheRouter catalogverified
Thinking modeSupported upstream for the qwen-plus family; OpenAI-compatible calls use enable_thinking upstream, while TheRouter exposes a reasoning parameterhelp.aliyun.com β†—verified
Function calling and toolsAlibaba's text-generation guide lists function calling for all general models and built-in tools for Plus-tier Qwen recommendationshelp.aliyun.com β†—verified
Structured outputSupported for qwen-plus family routes in Alibaba's capability tables and exposed by TheRouter as response_formathelp.aliyun.com β†—verified
ModalitiesText input β†’ text outputhelp.aliyun.com β†—verified

Benchmarks

BenchmarkDistributionScoreSource
Public Plus-tier benchmark table
Alibaba's text-generation guide recommends Plus-tier Qwen for balanced capability and cost, but the fetched public page does not publish a numeric benchmark table for the `qwen-plus` alias.
β€”Not publicly disclosedβ€”
Thinking-mode benchmark
The deep-thinking documentation explains how to enable thinking and lists qwen-plus-family support, but does not publish qwen-plus reasoning scores.
β€”Not publicly disclosedβ€”
TheRouter latency / throughput benchmark
No public TTFT or throughput measurement is published for TheRouter's `qwen/qwen-plus` route. Measure with thinking both on and off before setting production SLOs.
β€”Not publicly disclosedβ€”

API Usage Examples

Use the global api.therouter.ai endpoint shown below for new integrations; the legacy China accelerated endpoint is retired.

cURL
curl https://api.therouter.ai/v1/chat/completions   -H "Content-Type: application/json"   -H "Authorization: Bearer $THE_ROUTER_API_KEY"   -d '{
    "model": "qwen/qwen-plus",
    "messages": [
      {"role": "user", "content": "Summarize the key points from this input."}
    ]
  }'

Chat completion

Use the OpenAI-compatible chat endpoint for the default non-thinking path. Add reasoning only after validating how your SDK serialises vendor-specific options.

cURL
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen/qwen-plus","messages":[{"role":"user","content":"Summarize this customer feedback in three bullets."}],"max_tokens":600}'

More from qwen

Similar models

Cross-provider sibling models

News & changes

2026-08-02

Alibaba's text-generation guide positions Plus-tier Qwen as the balanced default

The current Model Studio guidance recommends Plus-tier Qwen models for office work, chatbots, content generation, document processing, and agent or coding workflows that need 1M context, tools, and balanced cost. The unsuffixed qwen-plus route remains listed as a legacy Qwen alias with the same broad capability columns.

re-authored by TheRouterhelp.aliyun.com β†—

Frequently asked

What is qwen/qwen-plus in TheRouter?

It is TheRouter's OpenAI-compatible route for Alibaba Bailian's commercial qwen-plus alias: a text-in/text-out model with 1M-class context, 32K maximum output, function calling, structured output, and optional reasoning support in the catalog.

Does qwen-plus support thinking mode?

Yes, Alibaba's deep-thinking documentation lists qwen-plus-family models among hybrid-thinking routes. Upstream examples use enable_thinking; TheRouter exposes reasoning, so test the exact request body your SDK sends before treating reasoning traces as a stable contract.

When should I choose qwen-plus instead of qwen-long?

Choose qwen-plus for normal production chat, agents, structured output, and document tasks that fit within 1M tokens. Choose qwen-long only when context size is the bottleneck and you accept a narrower long-document route.

Is qwen-plus a fixed Qwen3.7 checkpoint?

No. The public route is an unsuffixed commercial alias. Alibaba separately lists dated Plus snapshots such as Qwen3.7 Plus and older qwen-plus snapshots, so version-specific regressions or benchmark claims should be pinned to a snapshot route when available.

Fact ledger β€” every claim on this page traces here
sourceURLretrieved
Context lengthhelp.aliyun.com β†—2026-08-02verified
Maximum outputβ€”β€”verified
Thinking modehelp.aliyun.com β†—2026-08-02verified
Function calling and toolshelp.aliyun.com β†—2026-08-02verified
Structured outputhelp.aliyun.com β†—2026-08-02verified
Modalitieshelp.aliyun.com β†—2026-08-02verified
Public Plus-tier benchmark tablehelp.aliyun.com β†—2026-08-02unknown
Thinking-mode benchmarkhelp.aliyun.com β†—2026-08-02unknown
TheRouter latency / throughput benchmarkhelp.aliyun.com β†—2026-08-02unknown
Alibaba's text-generation guide positions Plus-tier Qwen as the balanced defaulthelp.aliyun.com β†—2026-08-02verified
What is qwen/qwen-plus in TheRouter?help.aliyun.com β†—2026-08-02to verify
Does qwen-plus support thinking mode?help.aliyun.com β†—2026-08-02to verify
When should I choose qwen-plus instead of qwen-long?help.aliyun.com β†—2026-08-02to verify
Is qwen-plus a fixed Qwen3.7 checkpoint?help.aliyun.com β†—2026-08-02to verify
Help & contact