Qwen3.8-27B Dense: Local Deployment Guide and API Routing for the Best Open-Weight Vision-Language Model
Qwen3.8-27B is a 27.78-billion-parameter dense vision-language model that runs on a single GPU. This guide covers local deployment with vLLM and SGLang, DashScope API access, hardware requirements, and how to route requests through an OpenAI-compatible gateway.
Qwen3.8-27B is a 27.78-billion-parameter dense vision-language model released by Alibaba's Qwen team on August 14, 2026. It accepts text, image, and video input, ships under Apache 2.0, and has a native 262,144-token context window. The key word is dense — unlike the 2.4-trillion-parameter Qwen3.8-Max (sparse MoE), this model fits on a single GPU, making it realistic for teams that want to self-host a strong multimodal model without a multi-node cluster.
We wrote this guide because the dense vs. MoE deployment decision is practical, not theoretical. If you can self-host Qwen3.8-27B for your coding, document analysis, or agentic workloads, you eliminate per-token API costs entirely. When local capacity runs out — peak hours, burst traffic, or workloads that exceed your GPU's context budget — you fall back to a managed API through an OpenAI-compatible routing layer.
Qwen3.8-27B at a Glance
| Spec | Value |
|---|---|
| Parameters | 27,781,427,952 (dense) |
| Architecture | 64 decoder layers (48 Gated DeltaNet + 16 full-attention) |
| Modality | Text + Image + Video input, Text output |
| Context window | 262,144 tokens native (up to 1M via YaRN) |
| License | Apache 2.0 |
| Thinking mode | Yes (thinking + non-thinking) |
| Hidden / FFN size | 5,120 / 17,408 |
| Attention | 24 query heads, 4 KV heads, head dim 256 (GQA) |
| DashScope model ID | qwen3.8-27b |
| Hugging Face | Qwen/Qwen3.8-27B |
Sources: Hugging Face model card (retrieved 2026-08-21), DashScope newly-released models (retrieved 2026-08-21), Kingy.ai specifications review (retrieved 2026-08-21).
Why Dense Matters for Self-Hosting
The Qwen3.8 family now has two deployment-relevant form factors:
| Model | Architecture | Total params | Activation params | Min VRAM (full precision) | Single-GPU viable? |
|---|---|---|---|---|---|
| Qwen3.8-Max | Sparse MoE | 2.4T | ~95B per step | Multi-node cluster | No |
| Qwen3.8-27B | Dense | 27.78B | 27.78B (all active) | ~56 GB (BF16) | Yes |
A 2.4-trillion-parameter MoE model needs multiple high-end GPUs just to hold the weight shards. Even inference-optimized clusters cost thousands of dollars per month in GPU rental. The 27B dense model, by contrast, fits in a single 80 GB GPU at full BF16 precision, or in a 24 GB consumer GPU at 4-bit quantization.
This is the practical case for Qwen3.8-27B: teams that cannot or do not want to provision a cluster for Qwen3.8-Max can run a dense model locally and still get strong coding, reasoning, and vision capabilities.
For a detailed comparison of the MoE and open-weight variants, see our Qwen3.8 Open-Source 2.4T vs. Proprietary Max Comparison.
Benchmark Performance
Published results from the Qwen team (all scores are Qwen-reported; independent reproductions are still emerging as of August 21, 2026):
| Benchmark | Qwen3.8-27B | Qwen3.6-27B | Delta |
|---|---|---|---|
| Terminal-Bench 2.1 | 73.0 | 63.4 | +9.6 |
| DeepSWE 1.1 | 42.2 | 13.3 | +28.9 |
| SWE-bench Pro | 61.7 | — | — |
| OSWorld-Verified | 84.3 | 63.9 | +20.4 |
| SWE-MM | 38.6 | 25.7 | +12.9 |
| GPQA Diamond | 89.2 | — | — |
| Humanity's Last Exam | 30.8 | — | — |
The coding and agentic improvements over Qwen3.6-27B are substantial. Terminal-Bench, DeepSWE, and OSWorld all show large jumps, which aligns with DashScope's release notes describing "improved coding and office scenario capabilities."
Caveat: Every benchmark score listed above comes from Qwen's own evaluation. Several benchmarks use in-house or modified harnesses. Independent evaluation from the community is still in early stages.
Sources: Hugging Face model card (retrieved 2026-08-21), Kingy.ai benchmark analysis (retrieved 2026-08-21), Northflank deployment guide (retrieved 2026-08-21).
Hardware Requirements
Full-Precision Serving (BF16 / FP8)
The model weights alone occupy approximately 55.6 GB in BF16. Add KV cache overhead for a reasonable context window, and you need:
- 80 GB GPU (A100/H100/H200): Comfortable for BF16 serving with moderate context.
- 48 GB GPU (A6000/L40S): Works with FP8 quantization and careful context budget.
Quantized Serving (4-bit / GGUF)
Third-party GGUF conversions are available on Hugging Face. At 4-bit quantization:
- 24 GB GPU (RTX 4090 / RTX 5090): Fits the 4-bit model with limited context (~8K-16K tokens practical).
- 32-48 GB unified memory (Apple M-series): Runs via Ollama or llama.cpp with room for longer context.
Important: The 24 GB number is for weights only. KV cache grows with context length. A 24 GB GPU running a 4-bit quant at 262K context will run out of memory. Budget 32-48 GB for comfortable quantized serving with meaningful context.
Production Recommendation
For production self-hosting with reasonable throughput and context:
# Single H100 80GB or A100 80GB
# Serves BF16 or FP8 with 32K-128K context comfortably
# Expect 40-80 tokens/second output depending on batch size
Sources: Kingy.ai hardware analysis (retrieved 2026-08-21), Northflank GPU requirements (retrieved 2026-08-21).
Local Deployment with vLLM
vLLM provides high-throughput serving with an OpenAI-compatible API out of the box.
pip install vllm
vllm serve Qwen/Qwen3.8-27B \
--host 0.0.0.0 \
--port 8000 \
--max-model-len 32768 \
--dtype bfloat16
Once the server is running, you can call it with any OpenAI SDK client:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="not-needed",
)
response = client.chat.completions.create(
model="Qwen/Qwen3.8-27B",
messages=[
{"role": "user", "content": "Explain the difference between dense and MoE model architectures in three sentences."}
],
)
print(response.choices[0].message.content)
For FP8 serving on an H100 to save VRAM:
vllm serve Qwen/Qwen3.8-27B \
--host 0.0.0.0 \
--port 8000 \
--max-model-len 65536 \
--dtype float16 \
--quantization fp8
Local Deployment with SGLang
SGLang is another high-performance option with native Qwen3.8-27B support:
pip install sglang[all]
python -m sglang.launch_server \
--model-path Qwen/Qwen3.8-27B \
--host 0.0.0.0 \
--port 8000 \
--context-length 32768
SGLang also exposes an OpenAI-compatible /v1/chat/completions endpoint, so the same client code from the vLLM section works without changes.
Source: SGLang Qwen3.8-27B cookbook (retrieved 2026-08-21).
Local Deployment with Ollama (Consumer Hardware)
For workstation or laptop use, Ollama provides the simplest path:
ollama run qwen3.8:27b
This pulls a quantized GGUF automatically and runs on Apple Silicon (32 GB+ recommended) or NVIDIA GPUs with 24 GB+ VRAM. The OpenAI-compatible API is available at http://localhost:11434/v1.
DashScope Managed API
If you do not want to self-host, Qwen3.8-27B is available on DashScope (Alibaba Cloud Model Studio) with model ID qwen3.8-27b.
Pricing (Beijing region)
As of August 21, 2026, the individual pricing for qwen3.8-27b on DashScope has not been separately listed in the pricing page. We expect it to be priced in the same tier as other 27B-class models. Check the DashScope pricing page for the latest rates.
Third-Party API Providers
Qwen3.8-27B is also available through third-party inference providers:
| Provider | Input (per 1M tokens) | Output (per 1M tokens) | Source |
|---|---|---|---|
| OpenRouter | $0.40 | $3.00 | openrouter.ai (retrieved 2026-08-21) |
| SiliconFlow | ~$0.30-0.45 | ~$3.20 | siliconflow.com/pricing (retrieved 2026-08-21) |
For teams already using SiliconFlow or DashScope through an OpenAI-compatible routing layer, adding Qwen3.8-27B is a configuration change — point the model ID and the requests route automatically.
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
DashScope API Quick Start
from openai import OpenAI
client = OpenAI(
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
api_key="sk-your-dashscope-key",
)
response = client.chat.completions.create(
model="qwen3.8-27b",
messages=[
{"role": "user", "content": "What are the advantages of dense models over MoE for local deployment?"}
],
)
print(response.choices[0].message.content)
Vision Input
Qwen3.8-27B is a native vision-language model. Pass images in the message content:
response = client.chat.completions.create(
model="qwen3.8-27b",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "Describe what you see in this architecture diagram."},
{"type": "image_url", "image_url": {"url": "https://example.com/diagram.png"}},
],
}
],
)
Hybrid Routing: Local-First with API Fallback
The most cost-effective deployment pattern for Qwen3.8-27B combines local self-hosting with managed API fallback. The idea is straightforward:
- Primary: Route requests to your local vLLM/SGLang instance.
- Fallback: When local capacity is exceeded (queue depth, GPU utilization, or context length beyond your VRAM budget), route to DashScope or SiliconFlow.
- Override: For workloads that need the full MoE flagship (long-context analysis, complex multi-step reasoning), route directly to
qwen3.8-maxon DashScope.
This pattern works because Qwen3.8-27B and the managed DashScope API both speak OpenAI-compatible chat completions. An OpenAI-compatible routing gateway can switch between local and cloud endpoints without changing application code.
TheRouter routes OpenAI-compatible requests through configured providers. If you configure a local endpoint as a provider alongside DashScope, the routing layer handles failover automatically.
When to Self-Host vs. When to Use the API
| Scenario | Recommendation |
|---|---|
| Consistent high-volume inference (thousands of requests/hour) | Self-host — amortized GPU cost beats per-token API pricing |
| Privacy-sensitive workloads (medical, legal, financial documents) | Self-host — data never leaves your infrastructure |
| Burst traffic or unpredictable load | API — pay per token, scale instantly |
| Need for 1M+ token context | API (Qwen3.8-Max) — local 27B with YaRN at 1M is experimental |
| Coding agent with moderate context (8K-32K) | Self-host — sweet spot for 27B dense on a single GPU |
| Multimodal document analysis | Either — depends on volume and latency requirements |
Common Deployment Issues
Out of Memory at Long Context
The 262K native context does not mean your GPU can handle 262K tokens. KV cache memory grows linearly with sequence length. On a 24 GB GPU with 4-bit weights, practical context is 8K-16K. Budget 80 GB for 128K+ context.
Slow First-Token Latency
Dense 27B models have higher first-token latency than smaller models like Qwen3.7-Flash. If you need sub-200ms first-token response, consider serving with FP8 on an H100 or using the managed API.
GGUF Quality vs. Official Weights
Only BF16 and FP8 checkpoints on Hugging Face are official Qwen artifacts. All GGUF files are third-party conversions. Quality at aggressive quantization levels (2-bit, 3-bit) degrades noticeably. We recommend 4-bit (Q4_K_M or Q4_K_S) as the practical floor for production use.
Production Checklist
- GPU sized correctly — 80 GB for BF16/FP8 production, 24-48 GB for quantized
- Context budget set —
--max-model-lenmatched to actual VRAM, not the 262K maximum - Health checks configured — vLLM and SGLang both expose
/healthendpoints - Fallback provider configured — DashScope or SiliconFlow as secondary when local is down
- Rate limits set — Protect your GPU from queue explosion during traffic spikes
- Monitoring in place — Track GPU utilization, queue depth, and p99 latency
Qwen3.8-27B vs. Comparable Open-Weight Models
| Model | Params | Dense/MoE | Vision | Context | License | SWE-bench Pro |
|---|---|---|---|---|---|---|
| Qwen3.8-27B | 27.78B | Dense | Yes | 262K | Apache 2.0 | 61.7 |
| Qwen3.6-27B | ~27B | Dense | Yes | 131K | Apache 2.0 | — |
| DeepSeek-V4-Flash | 284B total / 13B active | MoE | No | 1M | Model License | — |
| GLM-5.3 | Open-weight | Dense | No | 1M | — | — |
In the 27B-class dense model space, Qwen3.8-27B currently has no direct competitor that matches its combination of vision input, coding benchmarks, and Apache 2.0 licensing.
Summary
Qwen3.8-27B is the strongest open-weight dense model we have seen for local deployment of multimodal, coding, and agentic workloads. It runs on a single GPU, speaks OpenAI-compatible API, and can be combined with managed API fallback for a cost-effective hybrid architecture.
If you are evaluating self-hosted models for coding agents, document analysis, or privacy-sensitive inference, Qwen3.8-27B is the checkpoint to test. Start with the DashScope API to validate your use case, then move to local serving when the economics make sense.
Sources cited in this article:
- Hugging Face — Qwen/Qwen3.8-27B model card (retrieved 2026-08-21)
- DashScope — Newly Released Models (retrieved 2026-08-21)
- DashScope — Model Pricing (retrieved 2026-08-21)
- Kingy.ai — Qwen3.8-27B Specs, Benchmarks and Verdict (retrieved 2026-08-21)
- Kingy.ai — Qwen3.8-27B Local Hardware Guide (retrieved 2026-08-21)
- Northflank — Qwen3.8-27B Performance, Benchmarks, GPU Requirements (retrieved 2026-08-21)
- SGLang — Qwen3.8-27B Deployment Cookbook (retrieved 2026-08-21)
- OpenRouter — Qwen3.8-27B Pricing (retrieved 2026-08-21)
- SiliconFlow — Pricing (retrieved 2026-08-21)