← All articles

Kimi K3 vs Qwen3.8 Max vs Claude Opus 5: Frontier Reasoning Model API Comparison (August 2026)

A head-to-head API comparison of the three frontier reasoning models available in August 2026 — Moonshot's Kimi K3 (2.8T), Alibaba's Qwen3.8 Max (2.4T), and Anthropic's Claude Opus 5. Covers architecture, pricing, benchmarks, API surface, context windows, and when to pick each for production workloads.

· TheRouter

Three frontier reasoning models landed within six weeks of each other in the summer of 2026: Moonshot's Kimi K3 on July 16, Alibaba's Qwen3.8 Max on August 3, and Anthropic's Claude Opus 5 on July 24. All three ship 1M-token context windows, always-on reasoning, and OpenAI-compatible (or near-compatible) API surfaces. All three cost a fraction of what frontier models cost a year ago. And all three claim top-of-leaderboard performance on overlapping benchmark suites.

This comparison gives you the numbers, the architectural differences that produce those numbers, and a decision matrix for choosing between them. Every claim is attributed to a primary source. Where vendors report the same benchmark with different scores, we say so.

Sources: Kimi K3 Pricing, retrieved 2026-08-26; Kimi K3 Model Parameters, retrieved 2026-08-26; DashScope Model Pricing, retrieved 2026-08-26; Claude API Pricing, retrieved 2026-08-26; Qubrid K3 vs Qwen3.8 Max Comparison, retrieved 2026-08-26; Artificial Analysis, referenced via third-party reporting.

OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.

TL;DR Comparison Table

Kimi K3Qwen3.8 MaxClaude Opus 5
DeveloperMoonshot AIAlibaba (Qwen)Anthropic
API launchJuly 16, 2026August 3, 2026July 24, 2026
ArchitectureSparse MoESparse MoEDense transformer
Total parameters2.8T2.4TNot disclosed
Active params/token104B~95BNot disclosed
Context window1,048,5761,000,0001,000,000
Max outputNot published131K tokens128K tokens
Input price (per 1M)$3.00~$1.65 (¥12)$5.00
Output price (per 1M)$15.00~$4.95 (¥36)$25.00
Cache hit price$0.30/1M10% of input$0.50/1M
Reasoning controlreasoning_effortenable_thinkingExtended thinking
Open weightsYes (July 27)AnnouncedNo
Best forLong-context coding, SWEAgentic computer use, multimodalComplex reasoning, safety-critical

Architecture Comparison

Kimi K3: Hybrid Attention and Extreme Sparsity

K3's 2.8 trillion parameters make it the largest open-weight model available today. The architecture card, published alongside the weights on Hugging Face, details three efficiency mechanisms:

Kimi Delta Attention (KDA). Of the model's 93 layers, 69 use KDA and 24 use Gated Multi-head Latent Attention. This hybrid layout delivers roughly 6.3x faster decoding versus the K2 generation, per Kili Technology's August analysis. Moonshot shipped a vLLM implementation at launch, enabling day-0 self-hosting.

Stable LatentMoE. K3 routes 16 of 896 experts per token, plus 2 shared experts, yielding 104B activated parameters. The sparsity ratio of roughly 1:27 gives an approximately 2.5x scaling efficiency improvement over K2, per the model card.

Native MXFP4. K3 applies quantization-aware training from the SFT stage onward, with MXFP4 weights and MXFP8 activations. This means the 4-bit form is the trained form, not a post-hoc compression.

Qwen3.8 Max: Scale on the Qwen 3.5 Foundation

Alibaba's disclosure is architecturally thinner. What is confirmed: 2.4 trillion total parameters, roughly 95 billion active per token (sparsity ratio ~1:25), and native text/image/video input. The model builds on the Qwen 3.5 architectural foundation, per DataCamp's launch analysis.

Where Qwen3.8 Max differentiates is at the systems layer:

  • Dynamic Workflows for multi-agent orchestration with automatic plan decomposition
  • Vision-in-the-loop self-correction during agentic execution
  • Qwen-MM-Plugins for multimodal memory and visual tool use
  • Published rate limits of 2M TPM and 15K RPM

The architecture internals — layer count, expert count, routing scheme, attention mechanism — remain undisclosed as of this writing.

Claude Opus 5: Dense Frontier Reasoning

Anthropic's flagship takes the opposite architectural approach: a dense transformer with no publicly disclosed parameter count. Opus 5 launched on July 24, 2026, succeeding Opus 4.8 at the same $5/$25 price point. Key differences from the Chinese flagships:

  • Dense architecture rather than sparse MoE, trading parameter efficiency for training stability
  • Extended thinking with configurable max_tokens budget for reasoning
  • 1M-token context window (up from 200K in the 4.5 generation)
  • No open weights, with inference available only through Anthropic's API, AWS Bedrock, and Google Vertex AI

Benchmark Comparison

All benchmark figures below are vendor-reported unless marked otherwise. No single independent evaluator has published head-to-head results across all three models on the same benchmark version.

BenchmarkKimi K3Qwen3.8 MaxClaude Opus 5Source
Terminal-Bench 2.188.386.684.6 (Opus 4.8)Vendor tables, GMICloud
FrontierSWE81.273.5—Vendor tables, Qubrid comparison
OSWorld-Verified84.886.1—Vendor tables, Qubrid comparison
SWE-bench Pro—Trails Fable 5 by ~12 pts—emergent.sh
Intelligence Index (AA)5756—Artificial Analysis via The Decoder
Cost/Intelligence task$0.86$1.14—Artificial Analysis via The Decoder

Reading these numbers honestly:

  • K3 leads on terminal and repo-scale engineering tasks (Terminal-Bench, FrontierSWE)
  • Qwen3.8 Max leads on agentic computer-use tasks (OSWorld-Verified)
  • Claude Opus 5 benchmark data on these suites is incomplete at the time of writing. Opus 4.8 scored 84.6 on Terminal-Bench 2.1. Opus 5 likely matches or exceeds that, but we did not find published Opus 5 numbers on the same benchmarks
  • The Artificial Analysis Intelligence Index scores K3 narrowly above Qwen3.8 Max (57 vs 56), with K3 costing less per unit of intelligence ($0.86 vs $1.14 per task)

Pricing Comparison

When comparing API pricing across providers, always normalize to USD per million tokens and split input from output. Most providers price output tokens 2–5× higher than input tokens, so a workload heavy on completion length looks very different from a retrieval-heavy workload at the same nominal "price per million."

  • Use one currency (USD) — convert at publish date and cite the rate.
  • Split input/output — never quote a single blended number.
  • Cite each row to the provider's own pricing page with retrieval date.
  • Note context-window tiers — long-context pricing often steps higher.

Direct API Pricing (Per 1M Tokens)

InputOutputCache HitBatch
Kimi K3$3.00$15.00$0.30—
Qwen3.8 Max¥12 (~$1.65)¥36 (~$4.95)10% of input50% discount
Claude Opus 5$5.00$25.00$0.5050% discount

What the Pricing Means in Practice

For a 100K-token input with 2K-token output:

  • Kimi K3: $0.33 input + $0.03 output = $0.36
  • Qwen3.8 Max: $0.17 input + $0.01 output = $0.18
  • Claude Opus 5: $0.50 input + $0.05 output = $0.55

Qwen3.8 Max is roughly 3x cheaper than Opus 5 and 2x cheaper than K3 on a per-token basis. However, per-token cost is not the full picture. K3 uses fewer output tokens to reach the same quality on Artificial Analysis benchmarks ($0.86 vs $1.14 per intelligence task), which partially closes the per-token gap.

Both DashScope (Qwen3.8 Max) and Anthropic (Opus 5) offer batch processing at 50% discount. Moonshot does not currently offer batch pricing for K3.

API Surface Comparison

All three models support the OpenAI-compatible Chat Completions format, with provider-specific extensions:

Reasoning Control

# Kimi K3 — reasoning_effort (low / high / max)
response = client.chat.completions.create(
    model="kimi-k3",
    messages=[{"role": "user", "content": "Explain quantum entanglement"}],
    extra_body={"reasoning_effort": "high"}
)

# Qwen3.8 Max — enable_thinking (via DashScope)
response = client.chat.completions.create(
    model="qwen3.8-max",
    messages=[{"role": "user", "content": "Explain quantum entanglement"}],
    extra_body={"enable_thinking": True}
)

# Claude Opus 5 — extended thinking (Anthropic SDK)
response = anthropic_client.messages.create(
    model="claude-opus-5",
    max_tokens=16384,
    thinking={"type": "enabled", "budget_tokens": 10000},
    messages=[{"role": "user", "content": "Explain quantum entanglement"}]
)

Feature Support Matrix

FeatureKimi K3Qwen3.8 MaxClaude Opus 5
OpenAI Chat CompletionsYesYes (DashScope)Via proxy/adapter
StreamingYesYesYes
Tool callingYes (auto/none/required)YesYes
JSON modeYesYesYes
Vision inputYes (images)Yes (images, video)Yes (images)
Context cachingAutomaticExplicit + implicitTTL-based
Reasoning togglereasoning_effortenable_thinkingExtended thinking
Batch APINoYes (50% off)Yes (50% off)
temperature controlFixed at 1.0ConfigurableConfigurable

Context Window and Long-Document Handling

All three models support 1M-token context windows, but the details differ:

  • Kimi K3: 1,048,576 tokens total. Flat pricing across the entire context — no tiered rates for long inputs. Automatic context caching
  • Qwen3.8 Max: 1,000,000 tokens with a split: max input 991K, max input with thinking 983K, max output 131K, max reasoning budget 262K. DashScope supports both explicit and implicit caching with tiered pricing at 128K and 256K boundaries for some Qwen models (qwen3.8-max uses flat pricing up to 1M)
  • Claude Opus 5: 1,000,000 tokens. Max output 128K. TTL-based caching with 5-minute and 1-hour write tiers

For workloads that routinely push past 128K tokens, K3's flat pricing and Qwen3.8 Max's flat 1M-tier pricing both avoid the tiered surcharges that some older models charged.

Decision Matrix: When to Pick Each

Pick Kimi K3 When

  • Your workload is terminal-based coding and repo-scale software engineering — K3 leads on Terminal-Bench and FrontierSWE
  • You need open weights for self-hosting or compliance
  • You want the most intelligence per dollar based on Artificial Analysis benchmarks
  • Flat pricing across the full 1M context matters for your cost model
  • You prefer reasoning_effort levels (low/high/max) for granular reasoning control

Pick Qwen3.8 Max When

  • Your workload is agentic computer use — Qwen3.8 Max leads on OSWorld-Verified
  • Per-token cost is the primary constraint and you process high volumes
  • You need video input alongside text and images
  • You want batch processing at 50% discount
  • You are already on DashScope and want to stay within Alibaba's ecosystem
  • Multi-agent orchestration with Dynamic Workflows is part of your pipeline

Pick Claude Opus 5 When

  • Your workload requires the strongest general reasoning from an established safety-focused vendor
  • You need extended thinking with configurable token budgets for complex multi-step problems
  • Compliance and enterprise features (SOC 2, HIPAA BAA, AWS Bedrock / GCP Vertex) are requirements
  • You want Anthropic's safety and alignment guarantees for customer-facing applications
  • Batch processing at 50% discount is important and you can accept 24-hour turnaround

Route Through TheRouter When

You do not have to pick one. If your application routes OpenAI-compatible requests through TheRouter, you can set Kimi K3 as the primary for coding workloads, Qwen3.8 Max as the cost-optimized fallback, and Claude Opus 5 as the safety-critical path — all behind a single base_url. TheRouter supports model fallbacks that automatically retry on a secondary provider when the primary returns an error or times out.

FAQ

Is Kimi K3 available on DashScope? Yes. Kimi K3 landed on DashScope on August 19, 2026, making it accessible through the same OpenAI-compatible endpoint as Qwen models. DashScope pricing may differ from Moonshot's direct API pricing.

Are Qwen3.8 Max weights available for download? Alibaba announced open-weight release for the week of August 10, 2026. Check Hugging Face for the latest availability.

Which model has the best independent benchmark data? As of August 26, 2026, Artificial Analysis has published Intelligence Index scores for K3 and Qwen3.8 Max. Claude Opus 5 independent benchmark coverage on these specific suites is still building. All vendor-reported benchmarks should be treated as promotional until independently verified.

Can I use Claude Opus 5 through an OpenAI-compatible endpoint? Not directly from Anthropic's API, which uses a different request format (/v1/messages vs /v1/chat/completions). However, routing layers like TheRouter translate between the formats, letting you send OpenAI-compatible requests that reach Claude Opus 5 transparently.

How do the reasoning controls compare? K3 uses reasoning_effort with three levels (low, high, max). Qwen3.8 Max uses enable_thinking as a boolean toggle. Claude Opus 5 uses extended thinking with a budget_tokens parameter that lets you cap the reasoning token spend per request. K3's approach gives you coarse control, Opus 5's gives you fine-grained budget control, and Qwen3.8 Max's is a simple on/off switch.


This comparison reflects pricing, benchmark data, and feature availability as of August 26, 2026. Model capabilities and pricing change frequently; verify against the linked primary sources before making procurement decisions.

Models covered in this article

Help & contact