Back to Models

Kimi K2 Thinking

moonshotmoonshot/kimi-k2-thinking

API guide

Chat completion

Standard chat completion for K2 Thinking. The model will generate a thinking chain internally before producing its final answer. You can also stream the reasoning tokens.

cURL
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "moonshot/kimi-k2-thinking",
    "messages": [
      {"role": "system", "content": "You are a helpful AI assistant."},
      {"role": "user", "content": "Design a distributed caching strategy for a global e-commerce platform."}
    ],
    "max_tokens": 8192,
    "temperature": 0.7
  }'

Streaming with reasoning tokens

K2 Thinking exposes reasoning tokens in the streaming response. Use 'include_reasoning: true' in the request body to receive the model's chain-of-thought as a separate text delta before the final content.

cURL
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "moonshot/kimi-k2-thinking",
    "messages": [{"role": "user", "content": "Solve 3x^2 + 5x - 2 = 0"}],
    "stream": true,
    "include_reasoning": true
  }'

Tool calling (thinking + tools)

K2 Thinking natively interleaves reasoning tokens with tool calls during the chain-of-thought. The model can make 200–300 sequential tool invocations within a single reasoning trajectory, maintaining stable state throughout.

cURL
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "moonshot/kimi-k2-thinking",
    "messages": [
      {"role": "user", "content": "Find the latest research papers on MoE architectures and summarise the top 3 findings."}
    ],
    "tools": [{
      "type": "function",
      "function": {
        "name": "web_search",
        "description": "Search the web for information",
        "parameters": {
          "type": "object",
          "properties": {
            "query": {"type": "string", "description": "Search query"}
          },
          "required": ["query"]
        }
      }
    }],
    "tool_choice": "auto"
  }'
Fact ledger β€” every claim on this page traces here
sourceURLretrieved
Release datekimi-k2.org β†—2026-05-28verified
Architecturehuggingface.co β†—2026-05-28verified
Training costcnbc.com β†—2026-05-28to verify
Disk footprint (INT4)huggingface.co β†—2026-05-28verified
Licensehuggingface.co β†—2026-05-28verified
Supported inference engineshuggingface.co β†—2026-05-28verified
Training data cutoffβ€”β€”unknown
Humanity's Last Exam (w/ tools)lambda.ai β†—2026-05-28verified
BrowseComplambda.ai β†—2026-05-28verified
SWE-bench Verifiedlambda.ai β†—2026-05-28verified
AIME 2025 (w/ Python)www.reddit.com β†—2026-05-28to verify
GPQA-Diamondnist.gov β†—2026-05-28verified
MMLU-Pronist.gov β†—2026-05-28verified
CVE-Benchnist.gov β†—2026-05-28verified
SMT 2025nist.gov β†—2026-05-28verified
Moonshot AI releases Kimi K2 Thinking β€” open-weight reasoning model with 200+ sequential tool callskimi-k2.org β†—2026-05-28verified
NIST CAISI evaluates Kimi K2 Thinking β€” strongest PRC-developed open-weight model at time of releasenist.gov β†—2026-05-28verified
Kimi K2 Thinking running on 2Γ— M3 Ultra at 15 tok/s via mlx-lmsimonwillison.net β†—2026-05-28verified
What's the difference between Kimi K2 (Instruct) and Kimi K2 Thinking?kimi-k2.org β†—2026-05-28to verify
Can I run Kimi K2 Thinking on consumer hardware?lambda.ai β†—2026-05-28to verify
How does K2 Thinking compare to DeepSeek-R1?lambda.ai β†—2026-05-28to verify
Is K2 Thinking fully open source?huggingface.co β†—2026-05-28to verify
Help & contact