Kimi K3 vs Qwen3.8-Max vs DeepSeek V4 Pro: Chinese Frontier Model API Comparison for Coding Agents (September 2026)
A head-to-head comparison of three Chinese frontier model APIs for coding-agent workloads — Kimi K3, Qwen3.8-Max, and DeepSeek V4 Pro — covering architecture, pricing, coding benchmarks, tool calling, and routing configuration through a single OpenAI-compatible endpoint.
Three Chinese frontier models now compete head-to-head for coding-agent workloads: Kimi K3 (Moonshot AI, 2.8T dense parameters), Qwen3.8-Max (Alibaba, 2.4T MoE), and DeepSeek V4 Pro (DeepSeek, 1.6T MoE). All three support 1M-token context windows, built-in reasoning modes, and OpenAI-compatible endpoints. All three are routable through TheRouter today.
If you need the short answer: DeepSeek V4 Pro is the cheapest at $0.66–$1.32 per million input tokens, Kimi K3 is the densest and most expensive at $3/$15 per million input/output tokens, and Qwen3.8-Max scores highest on agent-oriented benchmarks like SWE-Bench Pro (67.7%) and CoWorkBench (74.8%). The right choice depends on whether you optimize for cost, raw coding capability, or agent reliability.
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
TL;DR Comparison Table
| Feature | Kimi K3 | Qwen3.8-Max | DeepSeek V4 Pro |
|---|---|---|---|
| Provider | Moonshot AI | Alibaba (DashScope) | DeepSeek |
| Parameters | 2.8T (dense) | 2.4T (MoE) | 1.6T (MoE) |
| Context window | 1M tokens | 1M tokens | 1M tokens |
| Max output | — | — | 384K tokens |
| Input price (cache miss) | $3.00/M | $2.00/M | $0.66–$1.32/M |
| Input price (cache hit) | $0.30/M | $0.25/M | $0.022–$0.044/M |
| Output price | $15.00/M | $6.00/M | $1.98–$3.96/M |
| Reasoning mode | Always-on (effort: max) | Thinking/non-thinking toggle | Thinking/non-thinking toggle |
| Vision | Yes | Text-only | Text-only (Flash has vision) |
| Tool calling | Yes | Yes | Yes |
| JSON mode | Yes | Yes | Yes |
| TheRouter model ID | moonshot/kimi-k3 | qwen/qwen3.8-max | deepseek/deepseek-v4-pro |
Pricing sources: Kimi K3 Pricing (retrieved 2026-09-18), DeepSeek Pricing (retrieved 2026-09-18), DashScope Pricing (retrieved 2026-09-18). DeepSeek V4 Pro prices show off-peak/peak range.
When comparing API pricing across providers, always normalize to USD per million tokens and split input from output. Most providers price output tokens 2–5× higher than input tokens, so a workload heavy on completion length looks very different from a retrieval-heavy workload at the same nominal "price per million."
- Use one currency (USD) — convert at publish date and cite the rate.
- Split input/output — never quote a single blended number.
- Cite each row to the provider's own pricing page with retrieval date.
- Note context-window tiers — long-context pricing often steps higher.
Architecture Showdown
Kimi K3 — The Dense 2.8T Giant
Kimi K3 is the only dense (non-MoE) model in this comparison. At 2.8 trillion parameters — all activated on every forward pass — K3 uses Moonshot AI's proprietary KDA (Key-Dimension Attention) mechanism to keep inference tractable despite the massive parameter count (source, retrieved 2026-09-18).
For coding agents, the dense architecture means K3 processes every token through its full representational capacity. That matters for complex multi-file refactors where the model needs to hold subtle dependency chains across an entire codebase in context. The trade-off is cost: at $15 per million output tokens, every generated line of code is expensive.
Qwen3.8-Max — Alibaba's 2.4T MoE Contender
Qwen3.8-Max uses a Mixture-of-Experts architecture with 2.4 trillion total parameters, activating a subset on each forward pass. Alibaba positions it as "a new bar for coding and cowork" — a model designed for multi-step agentic workflows where the model plans, writes, tests, and iterates autonomously (source, retrieved 2026-09-18).
The MoE design makes Qwen3.8-Max more cost-efficient per token than K3 while maintaining competitive benchmark performance. At $2/$6 per million input/output tokens, it sits in the middle of this three-way price bracket.
DeepSeek V4 Pro — The Cost Leader at 1.6T MoE
DeepSeek V4 Pro runs a 1.6 trillion parameter MoE architecture, the smallest of the three by total parameter count. What it lacks in raw scale, it makes up in cost efficiency: with off-peak pricing at $0.66/$1.98 per million input/output tokens, V4 Pro is roughly 7.5x cheaper than K3 on output and 3x cheaper than Qwen3.8-Max (source, retrieved 2026-09-18).
DeepSeek V4 Pro also offers the largest max output window in this comparison at 384K tokens — useful for agents that generate verbose debugging output or write entire modules in a single generation step.
Pricing Deep-Dive
Cost matters differently for coding agents than for chat. Agents generate far more output tokens than typical chat usage — a SWE-bench task might consume 50K–200K output tokens across plan-write-test-fix loops. Here is what each model costs at typical coding-agent volumes:
| Scenario | Kimi K3 | Qwen3.8-Max | DeepSeek V4 Pro (off-peak) |
|---|---|---|---|
| Light task (10K in / 20K out) | $0.33 | $0.14 | $0.046 |
| Medium task (50K in / 100K out) | $1.65 | $0.70 | $0.231 |
| Heavy refactor (200K in / 500K out) | $8.10 | $3.40 | $1.122 |
| Full-day agent session (1M in / 2M out) | $33.00 | $14.00 | $4.62 |
The math is straightforward: if you run an agent 8 hours a day hitting medium-complexity tasks, DeepSeek V4 Pro at off-peak rates costs roughly what one Kimi K3 task costs. For teams running multiple agents in parallel, that difference compounds fast.
Cache Economics
All three models offer cache-hit discounts, but the savings spread is enormous:
- DeepSeek V4 Pro: cache-hit input drops to $0.022/M (off-peak) — a 97% discount from cache-miss price
- Kimi K3: cache-hit input at $0.30/M — a 90% discount
- Qwen3.8-Max: cache-hit input at $0.25/M — an 87.5% discount
For coding agents that repeatedly feed the same codebase context (repo files, documentation, test suites), DeepSeek's cache pricing is an order of magnitude cheaper than either competitor. If your agent workflow hits cache frequently — and well-designed agents do — DeepSeek's effective cost drops even further below the headline numbers.
Coding Benchmarks Compared
Benchmark data comes from vendor reports and independent evaluations. Not all models have been tested on identical suites, which limits direct comparison on some metrics.
| Benchmark | Kimi K3 | Qwen3.8-Max | DeepSeek V4 Pro |
|---|---|---|---|
| SWE-bench Verified | — | — | 80.6% |
| SWE-bench Pro | — | 67.7% | 55.4% |
| Terminal-Bench 2.1 | — | 86.6% | 87.9% |
| LiveCodeBench (Vals) | — | 87.9% | 87.5% |
| OpenHarmony Bench | — | 60.8% | 59.0% |
| NL2Repo | — | 55.9% | 61.5% |
| DeepSWE | — | 56.6% | 62.7% |
| Arena Code WebDev | #1 | — | — |
Sources: BenchLM comparison (retrieved 2026-09-18), Artificial Analysis (retrieved 2026-09-18). Dash (—) means the model has not been independently evaluated on that benchmark.
The benchmark picture is mixed, which is exactly why routing matters:
- Qwen3.8-Max leads on SWE-bench Pro (67.7% vs 55.4%), the harder variant of SWE-bench that tests multi-step bug fixes
- DeepSeek V4 Pro leads on Terminal-Bench 2.1 (87.9% vs 86.6%) and real-world repository generation tasks like NL2Repo and DeepSWE
- Kimi K3 leads Arena Code WebDev, a community-voted leaderboard for frontend code generation
No single model dominates all coding benchmarks. That is the argument for routing — different tasks benefit from different models.
Tool Calling and Agent Capabilities
For coding agents, tool calling is the critical API feature. An agent that cannot reliably call tools — file read/write, terminal execution, search — is useless regardless of its raw coding ability.
Kimi K3
K3 supports standard OpenAI-compatible tool calling plus "dynamic tools" that can be defined and modified mid-conversation. Moonshot reports that K3 was specifically trained for tool-calling reliability in agentic workflows (source, retrieved 2026-09-18). K3's reasoning effort is currently fixed at max — there is no way to reduce thinking overhead for simple tool calls, which can inflate latency and output token count on straightforward operations.
Qwen3.8-Max
Qwen3.8-Max supports both thinking and non-thinking modes with an explicit toggle. For tool-calling-heavy agent loops, you can switch to non-thinking mode to reduce latency and cost on routine tool invocations, then switch back to thinking mode for complex planning steps. This flexibility gives operators more control over the cost/quality trade-off at each step of an agent workflow.
Alibaba's CoWorkBench score (74.8%) specifically measures multi-step agent collaboration — tasks where the model must coordinate tool calls, track state, and recover from errors (source, retrieved 2026-09-18).
DeepSeek V4 Pro
DeepSeek V4 Pro supports thinking/non-thinking toggle, tool calls, and JSON mode. DeepSeek also provides both OpenAI-format and Anthropic-format API endpoints, which can simplify integration for teams that use Claude-style tool calling patterns. The 384K max output window means agents can produce very long tool-call sequences without hitting output truncation — a practical advantage for agents that generate verbose debugging traces.
Context Window Behavior
All three models claim 1M-token context windows, but real-world behavior differs:
-
Kimi K3: 1,048,576 tokens. Moonshot's dense architecture means the entire context is processed through all 2.8T parameters, which theoretically preserves attention quality at extreme context lengths but increases per-token latency as context grows.
-
Qwen3.8-Max: 1M tokens. Alibaba's MoE routing means context quality may degrade at extreme lengths (past ~500K tokens) as expert selection becomes noisier. In practice, most coding-agent sessions stay well under 500K tokens per turn.
-
DeepSeek V4 Pro: 1M-token context, 384K max output. The output cap is the binding constraint for long-running agents: if your agent generates more than 384K tokens in a single response, the output is truncated. For multi-turn agent loops where each turn generates 10K–50K tokens, this is rarely a practical limit.
For coding agents, context length matters most for repo-level tasks — feeding an entire codebase into context for a full-repo refactor. All three models handle this equally at the spec level. Real-world latency and quality at extreme context lengths are harder to compare without controlled testing.
Latency and Throughput
Coding agents are latency-sensitive in their inner loop (plan → write → run tests → read output → iterate). First-token latency and tokens-per-second throughput directly affect total task completion time.
Published speed benchmarks vary by provider and load, but the general pattern holds:
- DeepSeek V4 Pro: Fastest first-token latency in thinking mode across most independent evaluations, benefiting from aggressive MoE routing that activates fewer parameters per token
- Qwen3.8-Max: Moderate latency, with non-thinking mode significantly faster than thinking mode — useful for tool-call-heavy loops
- Kimi K3: Slowest of the three due to dense architecture, but reasoning-effort control is limited (always
max), so there is no way to trade quality for speed on simple steps
For agents that run 50+ turns per task, first-token latency compounds. A 2-second per-turn advantage across 50 turns saves nearly 2 minutes per task — meaningful at scale.
Decision Matrix
Pick Kimi K3 if:
- Your workload is frontend-heavy (K3 leads Arena Code WebDev)
- You need vision input alongside coding (K3 is the only model with native vision in this comparison)
- Budget is secondary to raw output quality
- You value dense-architecture attention quality for complex dependency chains
Pick Qwen3.8-Max if:
- You want the best SWE-bench Pro score (67.7%) among these three
- You need flexible reasoning-mode control (thinking/non-thinking toggle)
- Multi-step agent collaboration is your primary use case (CoWorkBench 74.8%)
- You want open weights for potential self-hosting
Pick DeepSeek V4 Pro if:
- Cost is the primary constraint ($0.66/$1.98 per M at off-peak)
- You run agents at high volume and need cache-hit pricing below $0.05/M
- You need the largest max output window (384K tokens)
- Your agents benefit from Anthropic-format API compatibility
- You operate primarily during off-peak hours (half-price)
TheRouter Configuration
All three models are available through TheRouter with OpenAI-compatible routing. Here is a coding-agent fallback chain that prioritizes cost for routine work and quality for hard tasks:
# Cost-optimized default: DeepSeek V4 Pro
curl https://api.therouter.ai/v1/chat/completions \
-H "Authorization: Bearer $THEROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4-pro",
"messages": [{"role": "user", "content": "Refactor the auth module to use JWT tokens"}],
"tools": [...]
}'
# Quality-optimized: Qwen3.8-Max for hard tasks
curl https://api.therouter.ai/v1/chat/completions \
-H "Authorization: Bearer $THEROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen/qwen3.8-max",
"messages": [{"role": "user", "content": "Debug and fix the failing integration test suite"}],
"tools": [...]
}'
For agents that route dynamically, TheRouter supports model fallback chains. A practical pattern for coding agents:
- Primary:
deepseek/deepseek-v4-pro— cost-efficient for routine coding tasks - Fallback on error or timeout:
qwen/qwen3.8-max— stronger on complex multi-step tasks - Final fallback:
moonshot/kimi-k3— when you need the densest representation and vision
This gives you cost-optimized routing by default with automatic escalation to higher-capability (and higher-cost) models when the primary fails or times out.
FAQ
Which model is best for Cursor or Claude Code routing?
For Cursor and Claude Code users routing through TheRouter, DeepSeek V4 Pro offers the best cost-to-quality ratio for general coding work. If your IDE agent frequently handles complex refactors across large codebases, Qwen3.8-Max provides stronger multi-step planning. Note that Cursor has known limitations with custom API endpoint routing in sub-agents — verify your configuration routes all requests through TheRouter, not just the primary model.
Can I use these models with the OpenAI Python SDK?
Yes. All three providers expose OpenAI-compatible endpoints, and TheRouter unifies them behind a single base_url. Set base_url="https://api.therouter.ai/v1" and use the model IDs listed in the comparison table (moonshot/kimi-k3, qwen/qwen3.8-max, deepseek/deepseek-v4-pro).
How do off-peak DeepSeek prices work?
DeepSeek defines peak hours as 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday. All other hours are off-peak at half the peak rate. For teams in US Pacific or European time zones, most working hours fall in DeepSeek's off-peak window, making the effective cost even lower.
Is Qwen3.8-Max actually open weight?
Yes. Alibaba has released Qwen3.8-Max weights, making it the only open-weight model in this comparison. You can self-host via services like Together, DeepInfra, or SiliconFlow — though the 2.4T parameter count requires significant GPU infrastructure.
Why is Kimi K3 so much more expensive?
K3's dense architecture means every token passes through all 2.8T parameters, requiring substantially more compute per token than MoE models where only a fraction of parameters activate. The premium pricing reflects this higher per-token compute cost. Moonshot positions K3 as a quality ceiling for tasks where representational capacity matters more than throughput.
Do these models support streaming?
Yes. All three support server-sent events (SSE) streaming via the standard OpenAI-compatible stream: true parameter. TheRouter preserves streaming behavior transparently when routing to any of these providers.
Which model handles the longest effective context?
All three claim 1M tokens. In practice, DeepSeek V4 Pro's 384K max output is the binding constraint for very long generation tasks. For input-heavy workloads (large codebase context), all three perform comparably at context lengths under 500K tokens. Beyond that, independent evaluations are sparse and results vary by task type.
All pricing and benchmark data retrieved on September 18, 2026. Prices and benchmark scores may have changed since publication. Verify current pricing on each provider's official documentation before making purchasing decisions.