Chinese LLM API Providers in 2026 H2: DashScope, DeepSeek, Kimi, Zhipu, Volcengine, and SiliconFlow Compared
A practical H2 2026 comparison of the six major Chinese LLM API providers: DashScope (Bailian), DeepSeek, Moonshot (Kimi), Zhipu (Z.ai), Volcengine Ark, and SiliconFlow. We cover flagship models, per-token pricing after recent repricing events, OpenAI SDK compatibility, rate limits, model lifecycle risk, and when each provider fits.
The fastest answer: pick DashScope when you want the Qwen3.8 stack with third-party model hosting and China-region compliance, DeepSeek when cost efficiency on reasoning-class workloads matters most after the Aug 16 repricing, Kimi when long-horizon coding or 1M-context agent tasks justify the premium, Zhipu when you need open-weight GLM-5.3 or a Z.ai coding subscription, Volcengine Ark when you are already in ByteDance's ecosystem or need Doubao Seed 2.1, and SiliconFlow when you want a free-tier open-model aggregator with an OpenAI-compatible endpoint. The landscape shifted meaningfully in H2 2026, and this page captures the current state.
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
When comparing API pricing across providers, always normalize to USD per million tokens and split input from output. Most providers price output tokens 2–5× higher than input tokens, so a workload heavy on completion length looks very different from a retrieval-heavy workload at the same nominal "price per million."
- Use one currency (USD) — convert at publish date and cite the rate.
- Split input/output — never quote a single blended number.
- Cite each row to the provider's own pricing page with retrieval date.
- Note context-window tiers — long-context pricing often steps higher.
Sources: DashScope Model Pricing, retrieved 2026-08-24; DeepSeek Models & Pricing, retrieved 2026-08-24; Kimi K3 Pricing, retrieved 2026-08-24; SiliconFlow Pricing, retrieved 2026-08-24; Zhipu GLM-5.3 specs, retrieved 2026-08-24; Volcengine Ark Pricing, retrieved 2026-08-24.
Why a Fresh Comparison for H2 2026
The H1 2026 LLM providers comparison covered the landscape as of late July. Since then, every Chinese provider made moves that change the math:
- DashScope launched Qwen3.8-Max (¥12/¥36 per 1M tokens), expanded third-party model hosting (DeepSeek, Kimi K3, GLM on one endpoint), and announced a major Oct 10, 2026 sunset wave that retires 30+ legacy model IDs.
- DeepSeek shipped V4-Pro GA with Responses API support and raised peak-hour pricing on Aug 16, introducing time-of-day billing for the first time.
- Moonshot released Kimi K3 at $3/$15 per 1M tokens with a 1M-context window, positioning it as a premium coding-and-reasoning flagship.
- Zhipu launched GLM-5.3 (743B parameters, open-weight) on Aug 14, continuing the open-source strategy while keeping Z.ai subscription plans.
- Volcengine Ark released Doubao Seed 2.1 with Pro and Turbo tiers, plus the Seed-Evolving continuous-iteration model.
- SiliconFlow became the go-to aggregator for hosting open-weight models at community prices, including DeepSeek V4-Flash at $0.13/$0.28 per 1M tokens.
If you are choosing a Chinese LLM API provider today, this page is the one to read.
TL;DR Provider Matrix
| Provider | Endpoint | OpenAI SDK Compatible | Flagship Model | Input / Output (per 1M tokens) | Context | Lifecycle Risk |
|---|---|---|---|---|---|---|
| DashScope (Bailian) | dashscope.aliyuncs.com | Yes (OpenAI-compat mode) | qwen3.8-max | ¥12 / ¥36 | 1M | Medium (Oct 10 sunset wave) |
| DeepSeek | api.deepseek.com | Yes | deepseek-v4-pro | $0.66 / $1.98 (off-peak) | 1M | Low |
| Moonshot (Kimi) | api.moonshot.cn | Yes | kimi-k3 | $3.00 / $15.00 | 1M | Low |
| Zhipu (Z.ai) | open.bigmodel.cn | Partial | GLM-5.3 | Pricing TBA (5.2: $1.40 / $4.40) | 1M | Low |
| Volcengine Ark | ark.cn-beijing.volces.com | Yes (OpenAI-compat mode) | doubao-seed-2.1-pro | ¥6 / ¥30 | 256K | Low |
| SiliconFlow | api.siliconflow.cn | Yes | Aggregator (V4-Flash, K3, GLM-5.2) | Varies (V4-Flash: $0.13 / $0.28) | Varies | Low |
All prices retrieved 2026-08-24. Peak/off-peak pricing applies to DeepSeek. DashScope prices are in RMB for the Beijing region. SiliconFlow prices are in USD.
DashScope (Bailian): The Qwen Stack + Third-Party Aggregator
DashScope is Alibaba Cloud's model inference platform. It is the primary host for the Qwen family and has quietly become a Chinese LLM aggregator by hosting third-party models from DeepSeek, Moonshot, Zhipu, and MiniMax on a single API key.
Flagship models (Aug 2026):
- qwen3.8-max — ¥12 input / ¥36 output per 1M tokens, 1M context, thinking + non-thinking modes, batch calling at 50% discount, context caching with separate pricing
- qwen3.7-max — Same pricing as 3.8-max but currently at 50% promotional discount (¥6/¥18 effective)
- qwen3.7-plus — ¥2 input / ¥8 output (256K context), 20% promotional discount
What changed in H2:
- Qwen3.8-Max launched with the same pricing tier as 3.7-Max, adding multimodal and longer-context capabilities
- Third-party model hosting expanded to include DeepSeek V4-Flash/Pro, Kimi K3, and GLM models through the same DashScope endpoint
- Oct 10, 2026 sunset wave will retire 30+ legacy model IDs including qwen-turbo, qwen-vl-plus, qwen-audio-turbo, and older Qwen3 snapshots
OpenAI SDK compatibility: Full base_url + api_key swap works. DashScope uses dashscope.aliyuncs.com/compatible-mode/v1 as the OpenAI-compatible endpoint. Tiered rate limits based on account level, not publicly documented per-model.
Best for: Teams that want the full Qwen lineup plus third-party models on one invoice, China-region compliance with Beijing/Shanghai/Shenzhen deployment, and the largest free-trial token grants (1M tokens per model, 90-day expiry).
Watch out for: The Oct 10 sunset wave requires proactive migration. If you use any qwen-turbo, qwen-vl, or qwen-audio model IDs, read our Oct 2026 DashScope sunset migration guide now.
Sources: DashScope Model Pricing, retrieved 2026-08-24; DashScope Model Deprecation, retrieved 2026-08-24.
DeepSeek: Cost Leadership with Peak/Off-Peak Billing
DeepSeek remains the cost leader for reasoning-class workloads. The V4 generation shipped in two tiers with an industry-first peak/off-peak pricing model.
Flagship models (Aug 2026):
- deepseek-v4-pro — $0.66 input / $1.98 output per 1M tokens (off-peak), doubled at peak hours. Cache hit: $0.022 (off-peak). 1M context, 384K max output. Responses API + Anthropic API compatible.
- deepseek-v4-flash — $0.22 input / $0.66 output per 1M tokens (off-peak). Same cache and peak multiplier. 2,500 concurrency limit vs 500 for Pro.
What changed in H2:
- V4-Pro went GA on Aug 13, replacing V3.2 as the recommended flagship
- Aug 16 repricing introduced peak/off-peak billing. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday. Off-peak is everything else. Peak rates are 2x off-peak.
- Responses API launched alongside V4-Pro GA, giving DeepSeek feature parity with OpenAI's post-Assistants API surface
- Cache hit pricing dropped to $0.022 per 1M input tokens (off-peak) for V4-Pro
OpenAI SDK compatibility: Full drop-in swap. DeepSeek also supports an Anthropic-format endpoint at api.deepseek.com/anthropic.
Best for: High-volume reasoning and coding workloads where cost matters. Off-peak V4-Flash at $0.22/$0.66 is hard to beat for quality-per-dollar. Batch workloads that can shift to off-peak hours save 50%.
Watch out for: Peak-hour pricing doubles your bill if your traffic patterns hit the UTC 01:00-04:00 and 06:00-10:00 windows (which overlap with Asia business hours). The 500-concurrent-request limit on V4-Pro is tight for high-throughput pipelines.
Sources: DeepSeek Models & Pricing, retrieved 2026-08-24.
Moonshot (Kimi): Premium Reasoning at 1M Context
Moonshot's Kimi K3 is the most expensive model on this list and positions itself as a premium coding-and-reasoning agent.
Flagship model (Aug 2026):
- kimi-k3 — $3.00 input / $15.00 output per 1M tokens. Cache hit: $0.30. 1M-token context. Always-on reasoning with configurable effort (low/high/max). Tool calling, JSON mode, structured output.
What changed in H2:
- Kimi K3 launched with industry-leading benchmarks on long-horizon coding (SWE-bench Verified, Codeforces) and 1M-context agentic tasks
- K3 became available on DashScope and SiliconFlow as a third-party hosted model
- Reasoning effort configuration added (
low/high/max), letting you trade quality for cost on lighter tasks
OpenAI SDK compatibility: Full base_url swap at api.moonshot.cn/v1. Supports tool calling with tool_choice constraints and dynamically loaded tools (K3-specific features).
Best for: Long-context coding agents, complex multi-step reasoning tasks, and workloads where a single model call replaces multiple cheaper calls. At $15/1M output tokens, K3 only makes economic sense when the task complexity justifies the premium.
Watch out for: K3 output pricing is 7.5x DeepSeek V4-Pro off-peak and 2.5x DashScope Qwen3.8-Max. For routine completions, cheaper models deliver equivalent results.
Sources: Kimi K3 Pricing, retrieved 2026-08-24; Kimi K3 Quickstart, retrieved 2026-08-24.
Zhipu (Z.ai): Open-Weight GLM with Coding Subscriptions
Zhipu takes a dual-track approach: open-weight models (available on HuggingFace, SiliconFlow, and DashScope) plus premium API access and coding subscriptions through Z.ai.
Flagship models (Aug 2026):
- GLM-5.3 — 743B parameter model, launched Aug 14, 2026. Open-weight release on HuggingFace. API pricing on BigModel/Z.ai not yet published at time of writing. GLM-5.2 remains the reference for pricing: $1.40 input / $4.40 output per 1M tokens (Z.ai API), $1.30 / $4.09 on SiliconFlow.
- GLM-5.2 — $1.40 input / $4.40 output per 1M tokens via Z.ai. 1M context. Cache hit: $0.26. Available on SiliconFlow and DashScope.
What changed in H2:
- GLM-5.3 launched as a 743B open-weight model with improved long-horizon coding benchmarks
- Z.ai coding subscription repriced to $18/$72/$160 per month tiers
- GLM-5.3 API pricing on BigModel not yet published; SiliconFlow and DashScope hosting expected to follow
OpenAI SDK compatibility: Partial. Z.ai's API at open.bigmodel.cn/api/paas/v4 accepts OpenAI-format requests for chat completions, but some features (web search, knowledge retrieval) use Zhipu-specific parameters.
Best for: Teams that want open-weight model access (self-hosting GLM-5.3) or the Z.ai coding subscription for development workflows. The open-weight path means you can run GLM-5.3 on your own infrastructure with no per-token cost.
Watch out for: GLM-5.3 API pricing not published yet. BigModel pricing page is client-rendered and hard to scrape programmatically. SiliconFlow pricing for GLM-5.2 ($1.30/$4.09) is slightly cheaper than direct Z.ai access.
Sources: GLM-5.3 Specs, retrieved 2026-08-24; GLM Pricing, retrieved 2026-08-24; SiliconFlow Pricing, retrieved 2026-08-24.
Volcengine Ark: ByteDance Ecosystem and Doubao Seed 2.1
Volcengine Ark is ByteDance's model inference platform, home to the Doubao (Seed) family and a growing catalog of third-party models.
Flagship models (Aug 2026):
- doubao-seed-2.1-pro — ¥6 input / ¥30 output per 1M tokens. Flagship reasoning model.
- doubao-seed-2.1-turbo — ¥3 input / ¥15 output per 1M tokens. Lower latency, lower cost.
- doubao-seed-evolving — ¥6 input / ¥30 output per 1M tokens. Continuously iterated model with hallucination reduction focus.
What changed in H2:
- Seed 2.1 launched with Pro and Turbo tiers, replacing Seed 2.0 as the recommended flagship
- Seed-Evolving introduced as a continuously updated model that receives rolling improvements without version bumps
- Third-party model hosting expanded (DeepSeek, Qwen models available on Ark)
OpenAI SDK compatibility: Yes, via ark.cn-beijing.volces.com/api/v3 with OpenAI-format request bodies. Requires Volcengine account and endpoint creation through the console.
Best for: Teams already using ByteDance infrastructure (TikTok, Lark, Volcengine cloud). The Seed-Evolving model is interesting for production workloads that want continuous improvement without manual model version bumps.
Watch out for: Volcengine docs are partly client-rendered, making automated discovery harder. Pricing is model-by-model with no platform-wide published rate card. Regional availability is primarily China mainland.
Sources: Volcengine Ark Pricing, retrieved 2026-08-24; Huoshan Token API Pricing, retrieved 2026-08-24.
SiliconFlow: Open-Model Aggregator with Free Tier
SiliconFlow aggregates open-weight and third-party models on a single OpenAI-compatible endpoint. It is not a model developer but an inference platform that competes on price and breadth.
Key models and pricing (Aug 2026):
| Model | Input / Output (per 1M tokens) | Context |
|---|---|---|
| DeepSeek V4-Flash | $0.13 / $0.28 | 1M |
| DeepSeek V4-Pro | $1.50 / $3.14 | 1M |
| Kimi K3 | $3.00 / $15.00 | 1M |
| GLM-5.2 | $1.30 / $4.09 | 1M |
| Kimi K2.5 | $0.45 / $2.25 | 262K |
| Qwen3.6-27B | $0.30 / $3.20 | 262K |
What changed in H2:
- Added Kimi K3 and DeepSeek V4-Pro/Flash to the hosted model catalog
- Free tier provides $1 in credits at signup, with some open-weight models (e.g., Qwen3.5-9B) available at near-zero cost
- Community pricing often undercuts direct provider APIs by 5-15%
OpenAI SDK compatibility: Full drop-in swap at api.siliconflow.cn/v1. Standard API key authentication.
Best for: Cost-conscious teams that want to test multiple Chinese models on one API key. The free tier and low entry cost make SiliconFlow ideal for prototyping and evaluation. Also useful as a secondary routing target for fallback scenarios.
Watch out for: SiliconFlow is a reseller, not a model developer. Model availability depends on upstream partnerships. Rate limits scale with account tier (free tier has restrictive per-model RPM limits). No SLA guarantees comparable to first-party providers.
Sources: SiliconFlow Pricing, retrieved 2026-08-24.
Decision Matrix: Pick the Right Provider by Workload
| Workload | Recommended Provider | Why |
|---|---|---|
| High-volume reasoning (cost-sensitive) | DeepSeek V4-Flash (off-peak) | $0.22/$0.66 per 1M, schedule batch jobs for off-peak |
| Premium coding agents | Kimi K3 or DeepSeek V4-Pro | K3 for max quality, V4-Pro for 3x lower output cost |
| China compliance + Qwen stack | DashScope | Beijing/Shanghai deployment, largest model catalog |
| Open-weight self-hosting | Zhipu GLM-5.3 | 743B open-weight, run on your own GPUs |
| ByteDance ecosystem | Volcengine Ark | Seed 2.1 + Lark/TikTok integration path |
| Prototyping + evaluation | SiliconFlow | Free tier, one key for multiple models |
| Multi-model fallback | TheRouter + any combination | Route across providers with model fallbacks |
TheRouter Routing Note
TheRouter routes OpenAI-compatible requests through configured providers where live product paths support it. For Chinese providers, this means you can configure DashScope, DeepSeek, Moonshot, or SiliconFlow as routing targets and use model fallbacks to shift traffic between them based on availability, latency, or cost. Volcengine Ark and Zhipu BigModel work through their respective OpenAI-compatible endpoints.
We do not claim all models listed above are always available through TheRouter. Check the live models page and providers page for current routing support.
FAQ
Q: Which provider is cheapest for reasoning workloads?
DeepSeek V4-Flash at off-peak pricing ($0.22/$0.66 per 1M tokens) is the lowest-cost option for reasoning-capable models as of Aug 2026. DashScope's Qwen3.7-Max at the promotional 50% discount (¥6/¥18) is competitive but the discount has no published end date.
Q: Can I use one API key for multiple Chinese models?
DashScope and SiliconFlow both offer multi-model access on a single API key. DashScope hosts Qwen + DeepSeek + Kimi + GLM models. SiliconFlow hosts open-weight and select proprietary models. Volcengine Ark also hosts some third-party models but requires endpoint creation per model.
Q: What happens on Oct 10, 2026?
DashScope retires 30+ legacy model IDs including qwen-turbo, qwen-vl-plus, qwen-audio-turbo, and older snapshots. API calls to retired IDs will return errors. See our migration guide for the full list and replacement mappings.
Q: Is DeepSeek's peak pricing really 2x?
Yes. Peak hours (01:00-04:00 and 06:00-10:00 UTC, weekdays) charge double across all V4 models. For Asia-based teams, this overlaps with morning business hours (09:00-18:00 CST). See our DeepSeek repricing guide for scheduling strategies.
This comparison reflects pricing and model availability as of August 24, 2026. Prices change; verify against the official pricing pages linked above before making procurement decisions.
Related posts: