Back to Models

Qwen Long

qwenqwen/qwen-long

10M-token long-context Qwen β€” built for document-scale analysis.

How TheRouter serves this differently from the vendor

As the vendor operates it

Alibaba Bailian documents qwen-long as a CN-hosted long-context model with 10M-token context, 8,192-token output, text-in/text-out, direct HTTP guidance of 1M tokens, and file submission recommended above that length.

On TheRouter

TheRouter exposes qwen/qwen-long on the OpenAI-compatible chat path with a catalog context length of 10,485,760 tokens and CN-region Bailian serving. The page therefore treats 10M as the vendor contract and 10,485,760 as TheRouter's operational catalog value, and warns that direct long prompts need validation.

qwen-long is the long-context specialist of the Qwen family. With a 10M-token context window it ingests entire books, multi-hundred-page legal filings, whole codebases, or weeks of chat history in a single request, then answers, summarises, or extracts across the full input.

This model is deployed only on the mainland Bailian endpoint (bailian-cn). It does not have a Singapore failover β€” if bailian-cn is unavailable the model is unavailable. Plan for that in your retry strategy.

10M token context
Largest commercial context in the Qwen family. No chunking, no map-reduce needed for most document-scale jobs.
Document analysis
Strong at cross-section reasoning, full-text Q&A, contract review, and multi-file code summarisation.
Cheap per token
$0.12 input / $0.45 output per MTok β€” extreme value for the context window size.
CN deployment only
Lives on bailian-cn. No Singapore failover β€” region-pinned to mainland.
When to use
Whole-document analysis, contract / regulation review, repo-wide code summarisation, long transcript Q&A β€” anywhere chunking would lose cross-section context.
When not to use
Latency-critical chat or workloads that require Singapore (bailian-sg) availability β€” qwen-long is CN-only and tuned for throughput on long inputs, not for fast turn-by-turn responses.
Pricing: $0.12 input / $0.45 output per MTok. CN deployment only β€” no Singapore failover.
Context Length
10.5M
Max Output
8K
Input Priceper 1M tokens
$0.1296/ 1M tokens
Output Priceper 1M tokens
$0.486/ 1M tokens

Modalities

text→text

Pricing Breakdown

TypeRate
Input$0.1296 / 1M tokens
Output$0.486 / 1M tokens

Supported Parameters

temperaturemax_tokenstop_presponse_formatstop

Specifications

Context length10,000,000 tokens documented by Alibaba; TheRouter catalog exposes 10,485,760 tokenshelp.aliyun.com β†—verified
Direct HTTP input limit1,000,000 tokens; Alibaba recommends file submission above this lengthhelp.aliyun.com β†—verified
Maximum output8,192 tokenshelp.aliyun.com β†—verified
ModalitiesText input β†’ text outputhelp.aliyun.com β†—verified
Unsupported upstream featuresFunction calling, structured output, web search, prefix completion, context cache, batch inference, and fine-tuninghelp.aliyun.com β†—verified
Published mainland China quotaqwen-long: 1,200 RPM and 3,000,000 TPM in cn-beijinghelp.aliyun.com β†—verified

Benchmarks

BenchmarkDistributionScoreSource
Public long-context benchmark suite
Alibaba's qwen-long page describes target tasks and limits but does not publish a benchmark table for this model.
β€”Not publicly disclosedβ€”
Document QA benchmark
Use your own long-document acceptance set before replacing a chunked retrieval pipeline.
β€”Not publicly disclosedβ€”
Latency / throughput benchmark
The model is designed for very large contexts; measure TTFT and completion latency at your intended prompt size.
β€”Not publicly disclosedβ€”

API Usage Examples

Use the global api.therouter.ai endpoint shown below for new integrations; the legacy China accelerated endpoint is retired.

cURL
curl https://api.therouter.ai/v1/chat/completions   -H "Content-Type: application/json"   -H "Authorization: Bearer $THE_ROUTER_API_KEY"   -d '{
    "model": "qwen/qwen-long",
    "messages": [
      {"role": "user", "content": "Summarize the key points from this input."}
    ]
  }'

Long-document chat

Use the OpenAI-compatible chat endpoint for long text prompts. Keep direct HTTP requests under the upstream 1M-token guidance unless your integration has an approved file-submission path.

cURL
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen/qwen-long","messages":[{"role":"user","content":"Summarize the attached long document text: ..."}],"max_tokens":1200}'

More from qwen

Similar models

Cross-provider sibling models

News & changes

2025-01-25

Alibaba lists qwen-long snapshot and qwen-long-latest alongside the base route

The qwen-long documentation lists the stable qwen-long route, a dynamically updated qwen-long-latest, and the snapshot qwen-long-2025-01-25. All three carry the same 10M context and 8,192-token output headline, while the documented quota differs between stable/latest and the snapshot route.

re-authored by TheRouterhelp.aliyun.com β†—

Frequently asked

Is qwen/qwen-long really a 10M-token model?

Alibaba documents qwen-long with 10,000,000-token maximum input and context length. The same page also says direct HTTP submission supports 1M tokens and recommends file submission above that length. TheRouter's catalog exposes 10,485,760 tokens, so treat 10M as the vendor-published contract and validate the exact request shape you plan to run.

Does qwen-long support function calling or structured output?

Alibaba's qwen-long capability table says function calling and structured output are not supported, along with web search, prefix completion, context cache, batch inference, and model fine-tuning. Through TheRouter, use plain chat prompts and treat JSON output as best-effort instruction following.

When should I choose qwen-long instead of qwen-plus?

Choose qwen-long only when context size is the constraint: whole-document review, very long transcript Q&A, cross-section extraction, or archive-level summarisation. For normal chat, short classification, or latency-sensitive flows, qwen-plus will usually be simpler and faster.

Fact ledger β€” every claim on this page traces here
sourceURLretrieved
Context lengthhelp.aliyun.com β†—2026-08-01verified
Direct HTTP input limithelp.aliyun.com β†—2026-08-01verified
Maximum outputhelp.aliyun.com β†—2026-08-01verified
Modalitieshelp.aliyun.com β†—2026-08-01verified
Unsupported upstream featureshelp.aliyun.com β†—2026-08-01verified
Published mainland China quotahelp.aliyun.com β†—2026-08-01verified
Public long-context benchmark suitehelp.aliyun.com β†—2026-08-01unknown
Document QA benchmarkhelp.aliyun.com β†—2026-08-01unknown
Latency / throughput benchmarkhelp.aliyun.com β†—2026-08-01unknown
Alibaba lists qwen-long snapshot and qwen-long-latest alongside the base routehelp.aliyun.com β†—2026-08-01verified
Is qwen/qwen-long really a 10M-token model?help.aliyun.com β†—2026-08-01to verify
Does qwen-long support function calling or structured output?help.aliyun.com β†—2026-08-01to verify
When should I choose qwen-long instead of qwen-plus?help.aliyun.com β†—2026-08-01to verify
Help & contact