Using Kimi K3 & DeepSeek V4.1 Flash via OpenAI-Compatible API: Integration Guide for TheRouter
Both Kimi K3 and DeepSeek V4.1 Flash expose OpenAI-compatible endpoints. Change your base_url and model ID, keep your existing OpenAI SDK code, and route both through TheRouter for automatic fallback. This guide covers setup, code snippets, feature matrix, gotchas, and cost comparison.
Both Kimi K3 (Moonshot AI) and DeepSeek V4.1 Flash expose OpenAI-compatible chat-completion endpoints. If you already use the OpenAI Python or Node.js SDK, integrating either model is a two-line change: swap base_url and model. No new SDK, no new auth flow, no new response format. Your existing code — streaming, tool calls, structured output — keeps working.
We wrote this guide because developers searching for "kimi api openai compatible" or "deepseek api openai compatible" land on generic provider docs that cover one model at a time. This post puts both side by side, shows the exact code for each, maps their feature differences, and walks through routing both through TheRouter so a single endpoint handles fallback automatically.
Sources: Kimi K3 Quickstart, retrieved 2026-09-12; Kimi K3 Pricing, retrieved 2026-09-12; DeepSeek API Docs, retrieved 2026-09-12; DeepSeek Pricing, retrieved 2026-09-12; DeepSeek V4.1 Flash Announcement, retrieved 2026-09-12; Model Parameter Reference, retrieved 2026-09-12.
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
Why OpenAI-Compatible Matters for These Two Models
The OpenAI chat-completion format (/v1/chat/completions) has become the lingua franca for LLM APIs. When a provider adopts it, every tool built for OpenAI — SDKs, agent frameworks, IDE extensions, observability platforms — works out of the box.
Kimi K3 and DeepSeek V4.1 Flash both adopted this format natively. That matters because these are two of the strongest models available in September 2026, and they come from providers with significantly different pricing structures, rate limits, and geographic availability. Being able to swap between them with a config change — or route between them automatically through a gateway like TheRouter — gives you optionality without rewriting integration code.
Kimi K3 — Base URL, Auth, Model IDs, and Code
Provider: Moonshot AI
Base URL: https://api.moonshot.ai/v1
Auth: Bearer token (get your key at platform.kimi.ai/console/api-keys)
Model ID: kimi-k3
Context window: 1,048,576 tokens (1M)
Max output: 64,000 tokens
Kimi K3 is a 2.8-trillion-parameter MoE model with native vision, always-on reasoning, and a 1M-token context window. It was the first open-source model to reach the 3T-parameter class.
Python example
from openai import OpenAI
client = OpenAI(
api_key="YOUR_MOONSHOT_KEY",
base_url="https://api.moonshot.ai/v1",
)
response = client.chat.completions.create(
model="kimi-k3",
messages=[{"role": "user", "content": "Explain API gateways in one paragraph."}],
)
print(response.choices[0].message.content)
cURL example
curl https://api.moonshot.ai/v1/chat/completions \
-H "Authorization: Bearer $MOONSHOT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "kimi-k3",
"messages": [{"role": "user", "content": "Explain API gateways in one paragraph."}]
}'
K3-specific parameters
reasoning_effort:"low","high", or"max"(default"max"). K3 always reasons; this controls depth. The olderthinkingparameter from K2 is not supported.tool_choice: supports"auto","none", and"required"(K2 models did not support"required").temperature,top_p,n,presence_penalty,frequency_penalty: all fixed and cannot be modified. Do not pass them explicitly.
Source: Model Parameter Reference, retrieved 2026-09-12.
DeepSeek V4.1 Flash — Base URL, Auth, Model IDs, and Code
Provider: DeepSeek
Base URL: https://api.deepseek.com (OpenAI format) or https://api.deepseek.com/anthropic (Anthropic format)
Auth: Bearer token (get your key at platform.deepseek.com/api_keys)
Model ID: deepseek-flash (the canonical name; legacy deepseek-v4-flash still routes to V4.1 Flash)
Context window: 1,000,000 tokens (1M)
Max output: 384,000 tokens
DeepSeek V4.1 Flash is a 552B-parameter MoE model using a new Causal Encoder-Decoder architecture with 8B active input parameters and 16B active output parameters. It launched on September 10, 2026, replacing V4 Flash, and benchmarks ahead of the previous flagship V4 Pro.
Python example
from openai import OpenAI
client = OpenAI(
api_key="YOUR_DEEPSEEK_KEY",
base_url="https://api.deepseek.com",
)
response = client.chat.completions.create(
model="deepseek-flash",
messages=[{"role": "user", "content": "Explain API gateways in one paragraph."}],
)
print(response.choices[0].message.content)
cURL example
curl https://api.deepseek.com/chat/completions \
-H "Authorization: Bearer $DEEPSEEK_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-flash",
"messages": [{"role": "user", "content": "Explain API gateways in one paragraph."}]
}'
V4.1 Flash-specific parameters
thinking:{"type": "enabled"}(default) or{"type": "disabled"}. Both thinking and non-thinking modes are supported.reasoning_effort:"low","medium", or"high"— controls reasoning depth when thinking is enabled.temperature,top_p,frequency_penalty,presence_penalty: all adjustable (unlike K3).- Peak/off-peak pricing: peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday. Off-peak rates are 50% of peak.
Source: DeepSeek API Docs, retrieved 2026-09-12; V4.1 Flash Announcement, retrieved 2026-09-12.
Side-by-Side Feature Matrix
| Feature | Kimi K3 | DeepSeek V4.1 Flash |
|---|---|---|
| Base URL | https://api.moonshot.ai/v1 | https://api.deepseek.com |
| Model ID | kimi-k3 | deepseek-flash |
| Parameters | 2.8T MoE | 552B MoE (8B/16B active) |
| Context window | 1M tokens | 1M tokens |
| Max output | 64K tokens | 384K tokens |
| Streaming | Yes | Yes |
| Tool calling | Yes (required supported) | Yes |
Structured output (json_schema) | Yes (strict: true) | Yes |
| Vision | Yes (image + video) | Yes (image) |
| Thinking mode | Always on (reasoning_effort) | Toggleable (thinking param) |
| Reasoning effort levels | low / high / max | low / medium / high |
| Temperature control | Fixed (cannot modify) | Adjustable |
| Anthropic API format | No | Yes (/anthropic path) |
| FIM completion | No | Yes (non-thinking mode) |
| Responses API | No | Yes |
| Concurrency limit | Tier-based | 2,500 (Flash) |
Routing Both Through TheRouter — Config and Fallback Pattern
TheRouter routes OpenAI-compatible requests through configured providers. You can point your application at a single TheRouter endpoint and let it handle provider selection, fallback, and load balancing.
A typical configuration for Kimi K3 with DeepSeek V4.1 Flash as fallback:
# TheRouter config excerpt
routes:
- model: "kimi-k3"
provider: moonshot
fallback:
- model: "deepseek-flash"
provider: deepseek
Your application code stays identical to the OpenAI SDK examples above — just point base_url at your TheRouter instance instead of a provider endpoint. If K3 returns a 429 (rate limit) or 5xx (server error), TheRouter automatically retries with DeepSeek V4.1 Flash.
For cost-optimized routing, you might reverse the order: use DeepSeek V4.1 Flash as primary (lower cost per token) and fall back to K3 for workloads that need vision input with video or higher reasoning effort.
See Model Fallbacks for TheRouter fallback configuration details.
Common Gotchas
1. Reasoning parameter mismatch. K3 uses reasoning_effort (top-level field). DeepSeek uses both thinking (to toggle reasoning on/off) and reasoning_effort (to set depth). If you route between both, your middleware needs to translate: K3 ignores thinking, and DeepSeek ignores K3-style reasoning_effort without thinking enabled. TheRouter handles this translation for supported parameters.
2. Fixed vs. adjustable temperature. K3 locks temperature at 1.0 and rejects any other value. DeepSeek lets you set it freely. If your code passes temperature=0.7, it works on DeepSeek but fails on K3 with invalid_request_error. Either omit temperature when targeting both, or handle the error in your fallback logic.
3. Model ID naming. DeepSeek retired deepseek-v4-flash as a model name — it now routes to V4.1 Flash for compatibility, but the canonical name is deepseek-flash. Use the new name to avoid confusion when debugging.
4. Cache invalidation on K3. Switching reasoning_effort between turns invalidates the prefix cache and costs you a cache miss ($3/M input vs. $0.30/M on cache hit). Pick your effort level before the conversation starts and keep it consistent.
5. Peak pricing on DeepSeek. Peak hours (01:00-04:00 and 06:00-10:00 UTC, Mon-Fri) charge double the off-peak rate. If your workload is flexible, scheduling batch jobs outside peak hours cuts input costs by 50%.
6. Max output differences. K3 caps output at 64K tokens; V4.1 Flash allows up to 384K. If you need very long outputs (code generation, document drafting), V4.1 Flash is the better pick.
Cost Comparison Table
All prices per 1M tokens, in USD.
| Cost component | Kimi K3 | DeepSeek V4.1 Flash (off-peak) | DeepSeek V4.1 Flash (peak) |
|---|---|---|---|
| Input (cache hit) | $0.30 | $0.003 | $0.006 |
| Input (cache miss) | $3.00 | $0.15 | $0.30 |
| Output | $15.00 | $0.60 | $1.20 |
DeepSeek V4.1 Flash is roughly 20x cheaper than K3 on input and 12-25x cheaper on output, depending on peak timing. K3's cost premium buys you a larger parameter count (2.8T vs. 552B), native video understanding, and consistently high reasoning depth.
For cost-sensitive workloads where raw reasoning quality matters less than throughput, V4.1 Flash is the clear pick. For tasks that need frontier-level reasoning, native video input, or structured output with K3-specific strict schemas, the cost difference may be justified.
Sources: Kimi K3 Pricing, retrieved 2026-09-12; DeepSeek Pricing, retrieved 2026-09-12.
Production Checklist
- Store API keys in environment variables or a secrets manager, not in code
- Set
max_tokensormax_completion_tokensto prevent runaway output costs - Handle
429rate-limit responses with exponential backoff, or use TheRouter's built-in retry - For K3: do not pass
temperature,top_p,n,presence_penalty, orfrequency_penalty - For DeepSeek: use
deepseek-flashas model ID (not the retireddeepseek-v4-flash) - For K3: decide
reasoning_effortbefore the conversation starts to preserve cache hits - For DeepSeek: schedule batch workloads outside peak hours (01:00-04:00, 06:00-10:00 UTC Mon-Fri) to halve input costs
- Test streaming behavior — both providers send
reasoning_contentdeltas beforecontentdeltas - If using tool calling with K3, consider the dynamic tool loading pattern to avoid filling the context window
- Monitor token usage per provider when routing through TheRouter
FAQ
Can I use the same OpenAI SDK code for both providers?
Yes. Both Kimi K3 and DeepSeek V4.1 Flash implement the /v1/chat/completions endpoint with the same request/response schema. Change base_url and model, and your code works.
Which model is better for coding tasks? Both are strong. K3 excels at long-horizon coding with visual feedback (game dev, frontend). V4.1 Flash benchmarks ahead of V4 Pro on coding tasks while being cheaper and faster. For pure code generation without vision input, V4.1 Flash offers better cost efficiency.
Does DeepSeek support the Anthropic API format too?
Yes. DeepSeek exposes an Anthropic-compatible endpoint at https://api.deepseek.com/anthropic. Kimi does not offer an Anthropic-format endpoint.
What happens if I pass temperature=0 to K3?
K3 rejects it with an invalid_request_error. K3 fixes temperature at 1.0 and does not accept modifications. Omit the parameter entirely.
Can TheRouter translate reasoning parameters between providers?
TheRouter routes OpenAI-compatible requests through configured providers and supports provider/model routing and fallback when live product paths support it. For provider-specific parameters like reasoning_effort or thinking, refer to TheRouter's parameter mapping documentation.
Is K3 available on DashScope (Alibaba Cloud)?
As of September 2026, Kimi K3 is available through Moonshot's own API (api.moonshot.ai) and through DashScope. Check the DashScope newly released models page for current availability.
What is the minimum top-up for K3? K3 is a flagship model that requires a minimum $1 top-up to unlock. Your cumulative top-up amount determines your rate limit tier.