Back to Models
Kimi K2 Thinking
moonshotmoonshot/kimi-k2-thinking
API guide
Chat completion
Standard chat completion for K2 Thinking. The model will generate a thinking chain internally before producing its final answer. You can also stream the reasoning tokens.
cURL
curl https://api.therouter.ai/v1/chat/completions \
-H "Authorization: Bearer $THEROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "moonshot/kimi-k2-thinking",
"messages": [
{"role": "system", "content": "You are a helpful AI assistant."},
{"role": "user", "content": "Design a distributed caching strategy for a global e-commerce platform."}
],
"max_tokens": 8192,
"temperature": 0.7
}'Streaming with reasoning tokens
K2 Thinking exposes reasoning tokens in the streaming response. Use 'include_reasoning: true' in the request body to receive the model's chain-of-thought as a separate text delta before the final content.
cURL
curl https://api.therouter.ai/v1/chat/completions \
-H "Authorization: Bearer $THEROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "moonshot/kimi-k2-thinking",
"messages": [{"role": "user", "content": "Solve 3x^2 + 5x - 2 = 0"}],
"stream": true,
"include_reasoning": true
}'Tool calling (thinking + tools)
K2 Thinking natively interleaves reasoning tokens with tool calls during the chain-of-thought. The model can make 200β300 sequential tool invocations within a single reasoning trajectory, maintaining stable state throughout.
cURL
curl https://api.therouter.ai/v1/chat/completions \
-H "Authorization: Bearer $THEROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "moonshot/kimi-k2-thinking",
"messages": [
{"role": "user", "content": "Find the latest research papers on MoE architectures and summarise the top 3 findings."}
],
"tools": [{
"type": "function",
"function": {
"name": "web_search",
"description": "Search the web for information",
"parameters": {
"type": "object",
"properties": {
"query": {"type": "string", "description": "Search query"}
},
"required": ["query"]
}
}
}],
"tool_choice": "auto"
}'Fact ledger β every claim on this page traces here
| source | URL | retrieved | |
|---|---|---|---|
| Release date | kimi-k2.org β | 2026-05-28 | verified |
| Architecture | huggingface.co β | 2026-05-28 | verified |
| Training cost | cnbc.com β | 2026-05-28 | to verify |
| Disk footprint (INT4) | huggingface.co β | 2026-05-28 | verified |
| License | huggingface.co β | 2026-05-28 | verified |
| Supported inference engines | huggingface.co β | 2026-05-28 | verified |
| Training data cutoff | β | β | unknown |
| Humanity's Last Exam (w/ tools) | lambda.ai β | 2026-05-28 | verified |
| BrowseComp | lambda.ai β | 2026-05-28 | verified |
| SWE-bench Verified | lambda.ai β | 2026-05-28 | verified |
| AIME 2025 (w/ Python) | www.reddit.com β | 2026-05-28 | to verify |
| GPQA-Diamond | nist.gov β | 2026-05-28 | verified |
| MMLU-Pro | nist.gov β | 2026-05-28 | verified |
| CVE-Bench | nist.gov β | 2026-05-28 | verified |
| SMT 2025 | nist.gov β | 2026-05-28 | verified |
| Moonshot AI releases Kimi K2 Thinking β open-weight reasoning model with 200+ sequential tool calls | kimi-k2.org β | 2026-05-28 | verified |
| NIST CAISI evaluates Kimi K2 Thinking β strongest PRC-developed open-weight model at time of release | nist.gov β | 2026-05-28 | verified |
| Kimi K2 Thinking running on 2Γ M3 Ultra at 15 tok/s via mlx-lm | simonwillison.net β | 2026-05-28 | verified |
| What's the difference between Kimi K2 (Instruct) and Kimi K2 Thinking? | kimi-k2.org β | 2026-05-28 | to verify |
| Can I run Kimi K2 Thinking on consumer hardware? | lambda.ai β | 2026-05-28 | to verify |
| How does K2 Thinking compare to DeepSeek-R1? | lambda.ai β | 2026-05-28 | to verify |
| Is K2 Thinking fully open source? | huggingface.co β | 2026-05-28 | to verify |