Qwen3.8 Open-Source (2.4T-A95B) vs Qwen3.8-Max: What the Open Weights Mean for API Routing
Alibaba released the open-weight Qwen3.8-2.4T-A95B on August 12, 2026 — the same 2.4-trillion-parameter MoE architecture behind the proprietary Qwen3.8-Max. This comparison covers the capability gaps, pricing across DashScope and SiliconFlow, and how to route between them through an OpenAI-compatible gateway.
Alibaba released Qwen3.8-Max as a proprietary API model on August 3, 2026. Nine days later, on August 12, the team dropped the open-weight checkpoint Qwen3.8-2.4T-A95B on Hugging Face and ModelScope. Both models share the same 2.4-trillion-parameter sparse Mixture-of-Experts architecture with 95 billion activated parameters per forward pass.
They are not the same product. The open-weight release ships text-only. The proprietary Max retains native vision and video input, structured outputs, built-in tools, and Alibaba's proprietary post-training. The pricing gap between them is equally large: third-party inference providers serve the open-weight checkpoint at a fraction of the Max's DashScope rate.
This comparison covers what changed between the two, where each one fits, and how to set up routing between them.
Architecture: Same Skeleton, Different Muscles
Both Qwen3.8-2.4T-A95B and Qwen3.8-Max are 2.4T-parameter MoE models that activate 95B parameters per token. Both accept a 1M-token context window. Both support thinking mode (extended chain-of-thought reasoning).
The open-weight release is a text-generation model only. It does not include the vision encoder. You cannot pass images or video frames to it. The proprietary Max retains full multimodal input — text, image, and video — with the same 1M context window.
The Max also ships with capabilities that depend on Alibaba's serving infrastructure: function calling, structured outputs (JSON mode), batch processing, prefix completion, and five built-in tools on the Responses API (code_interpreter, web_search, web_extractor, t2i_search, i2i_search). The open-weight checkpoint exposes none of these natively — though inference providers and frameworks like vLLM can add function calling and JSON mode on top.
| Feature | Qwen3.8-2.4T-A95B (open) | Qwen3.8-Max (proprietary) |
|---|---|---|
| Total parameters | 2.4T | 2.4T |
| Activated parameters | 95B | 95B |
| Context window | 1M tokens | 1M tokens |
| Input modalities | Text only | Text, image, video |
| Thinking mode | Yes | Yes |
| Function calling | Framework-dependent | Native |
| Structured output | Framework-dependent | Native (JSON mode) |
| Batch API | No | Yes |
| Built-in tools | No | 5 tools on Responses API |
| License | Apache 2.0 | Proprietary (API access) |
| Weights available | Yes (HuggingFace, ModelScope) | No |
Source: Qwen blog, ModelScope model card, MarktechPost. Retrieved 2026-08-20.
Pricing: The Cost Gap Is Substantial
Qwen3.8-Max on DashScope (Beijing region) costs ¥12 per million input tokens and ¥36 per million output tokens — roughly $2.00 and $6.00 at current exchange rates. Cached input drops to $0.25/M. There is no tiered pricing; the flat rate applies across the full 1M context window.
The open-weight Qwen3.8-2.4T-A95B is available through third-party inference providers at dramatically lower rates. SiliconFlow lists the model at $0.14 per million input tokens — a 93% reduction from the Max's DashScope rate. OpenRouter also hosts the open-weight checkpoint.
DashScope itself lists qwen3.8-2.4t-a95b as a model ID at the same ¥12/¥36 pricing as the Max. This is important: calling the open-weight model through DashScope costs the same as the Max. The cost advantage only materializes when you route through a third-party provider that serves the open weights on its own infrastructure.
| Provider | Model | Input (/M tokens) | Output (/M tokens) |
|---|---|---|---|
| DashScope | qwen3.8-max | $2.00 | $6.00 |
| DashScope | qwen3.8-2.4t-a95b | $2.00 | $6.00 |
| SiliconFlow | Qwen3.8-2.4T-A95B | $0.14 | varies |
| OpenRouter | qwen/qwen3.8-2.4t-a95b | varies | varies |
Sources: DashScope pricing, SiliconFlow (announcement post). Retrieved 2026-08-20.
When comparing API pricing across providers, always normalize to USD per million tokens and split input from output. Most providers price output tokens 2–5× higher than input tokens, so a workload heavy on completion length looks very different from a retrieval-heavy workload at the same nominal "price per million."
- Use one currency (USD) — convert at publish date and cite the rate.
- Split input/output — never quote a single blended number.
- Cite each row to the provider's own pricing page with retrieval date.
- Note context-window tiers — long-context pricing often steps higher.
When to Use Which
The decision breaks down along three axes: do you need vision, do you need the cheapest text inference, and do you need the proprietary tooling.
Choose Qwen3.8-Max when:
- Your workload includes image or video input (document OCR, screenshot analysis, video indexing)
- You need native function calling, structured outputs, or batch processing without framework overhead
- You want the built-in tools (code_interpreter, web_search) without managing them yourself
- DashScope's context caching ($0.25/M) makes Max cost-competitive for your prompt reuse pattern
Choose Qwen3.8-2.4T-A95B (open weights) when:
- Your workload is text-only — coding agents, research synthesis, document generation
- You want the lowest possible per-token cost and can route through SiliconFlow or self-host
- You need to keep data on your own infrastructure (compliance, air-gapped environments)
- You want to fine-tune the model for a specific domain
Route between both when:
- Most of your traffic is text-only (route to SiliconFlow for cost savings), but a minority of requests include images (route those to DashScope Max)
- You want a fallback chain: try the cheap open-weight endpoint first, fall back to Max if the provider is down
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
Routing Setup
Both models speak the OpenAI chat completions protocol. If you are using an OpenAI-compatible gateway, routing between them is a model ID and base URL change.
# Text-only request → route to SiliconFlow (open weights, lower cost)
from openai import OpenAI
client = OpenAI(
base_url="https://api.siliconflow.cn/v1",
api_key="YOUR_SILICONFLOW_KEY",
)
response = client.chat.completions.create(
model="Qwen/Qwen3.8-2.4T-A95B",
messages=[{"role": "user", "content": "Refactor this function to use async/await."}],
)
# Vision request → route to DashScope (proprietary Max, multimodal)
client = OpenAI(
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
api_key="YOUR_DASHSCOPE_KEY",
)
response = client.chat.completions.create(
model="qwen3.8-max",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "Describe what you see in this screenshot."},
{"type": "image_url", "image_url": {"url": "https://example.com/screenshot.png"}},
],
}],
)
An OpenAI-compatible routing layer can inspect the request content type and dispatch accordingly — text-only requests to the cheaper open-weight endpoint, multimodal requests to DashScope Max.
Benchmarks: How Close Is the Open-Weight Version?
Alibaba published benchmarks for Qwen3.8-Max but has not released a separate benchmark table for the open-weight checkpoint. The model card on ModelScope confirms the same architecture (2.4T total, 95B activated) but does not include evaluation scores.
Third-party evaluators on BenchLM scored Qwen3.8-Max at 79.9/100, ranking it sixth out of 221 models. Key scores include Terminal-Bench 2.1 at 86.6 (behind GPT-5.6 Sol's 88.8, ahead of Claude Opus 4.8's 84.6), GPQA Diamond at 92.6, and PaperBench at 93.0.
For the open-weight version, the expectation is near-parity on text-only benchmarks — the architecture and parameter count are identical. The gap will appear on any evaluation that requires vision input, since the open-weight release cannot process images.
Independent benchmarks specifically targeting Qwen3.8-2.4T-A95B are still emerging as of August 20, 2026.
Sources: BenchLM, Qwen blog, MarktechPost. Retrieved 2026-08-20.
Self-Hosting Considerations
The open-weight checkpoint is Apache 2.0 licensed. You can download it from Hugging Face or ModelScope and serve it on your own infrastructure.
The practical challenge is scale. At 2.4T total parameters, even with MoE sparsity (95B activated), serving this model requires substantial GPU memory. A single node with 8x H100 (80GB each, 640GB total) may not be sufficient depending on your serving framework's memory overhead for expert routing. Multi-node setups are the realistic path for production self-hosting.
For teams that want on-premise inference but cannot justify multi-node deployments, the Qwen3.8-27B dense model (released August 19, also Apache 2.0) fits on a single 80GB GPU and provides Qwen3.8-generation capabilities at a much smaller scale.
The practical self-hosting path for most teams: use SiliconFlow's hosted inference for the 2.4T open-weight model, and self-host the 27B dense variant for latency-sensitive or air-gapped workloads.
Summary
| Dimension | Open weights (2.4T-A95B) | Proprietary (Max) |
|---|---|---|
| Best price (text) | ~$0.14/M input via SiliconFlow | $2.00/M input via DashScope |
| Vision support | No | Yes |
| Tooling | Framework-dependent | Native |
| Self-hostable | Yes (Apache 2.0) | No |
| Routing fit | Default for text-only, high-volume | Multimodal, tooling-dependent |
The open weights turn Qwen3.8 from a single vendor's API into a multi-provider commodity. For text-only workloads, routing through SiliconFlow at $0.14/M versus DashScope at $2.00/M is a 93% cost reduction on the same architecture. For anything that touches images or video, the proprietary Max remains the only option in the Qwen3.8 family.
Setting up routing between the two is the pragmatic answer for most teams.
Sources cited in this post:
- Qwen3.8-Max announcement — Qwen blog (retrieved 2026-08-20)
- DashScope model pricing — Alibaba Cloud (retrieved 2026-08-20)
- Qwen3.8-2.4T-A95B model card — ModelScope (retrieved 2026-08-20)
- SiliconFlow pricing announcement — X/Twitter (retrieved 2026-08-20)
- Qwen3.8-Max benchmarks — BenchLM (retrieved 2026-08-20)
- MarktechPost Qwen3.8-Max coverage (retrieved 2026-08-20)
- Qwen API pricing analysis — Spheron (retrieved 2026-08-20)
- SiliconFlow pricing page (retrieved 2026-08-20)