DeepSeek V4.1 Flash API Complete Guide: 552B MoE Architecture, Native Vision, and Cost-Optimized Routing
Everything you need to build with DeepSeek V4.1 Flash — the 552B MoE model that activates just 8B parameters on input. We cover the Causal Encoder-Decoder architecture, native vision support, pricing (from $0.15/M input off-peak), thinking mode, tool calls, and how to route V4.1 Flash as a primary or fallback provider.
DeepSeek V4.1 Flash is a 552B-parameter Mixture-of-Experts model that activates only 8B parameters per input token and 16B per output token. It launched on September 10, 2026, replacing both V4 Flash and V4-Flash-Vision-Exp with a single multimodal endpoint. You call it by setting model: "deepseek-flash" against the same https://api.deepseek.com base URL — your existing OpenAI SDK code works without changes.
We wrote this guide because V4.1 Flash reshuffles DeepSeek's entire model lineup. V4 Flash and V4-Flash-Vision-Exp are retired immediately; V4 Pro requests will route to V4.1 Flash starting September 14. If you are running DeepSeek in production, you need to understand what changed, what the new pricing looks like, and where V4.1 Flash fits in a multi-provider routing setup.
Getting Started in 3 Minutes
Step 1: Create an account. Go to platform.deepseek.com and sign up with email or Google OAuth.
Step 2: Get your API key. Navigate to API Keys, click "Create API Key," and copy it. Keys are shown once.
Step 3: Make your first call.
pip install --upgrade openai
from openai import OpenAI
client = OpenAI(
api_key="YOUR_DEEPSEEK_KEY",
base_url="https://api.deepseek.com",
)
response = client.chat.completions.create(
model="deepseek-flash",
messages=[{"role": "user", "content": "What is DeepSeek V4.1 Flash?"}],
)
print(response.choices[0].message.content)
That is the entire integration. DeepSeek's API follows the OpenAI chat completions format, so any library or framework that supports OpenAI-compatible endpoints works out of the box.
OpenAI-compatible means a provider exposes a chat-completions endpoint whose request and response shape matches the OpenAI API contract closely enough that an unmodified OpenAI SDK call works against it after swapping three values: API key, base URL, and model name. The minimum surface in practice is POST /v1/chat/completions with messages, model, and an OpenAI-shaped streaming response.
Architecture Overview: Causal Encoder-Decoder with Asymmetric Activation
V4.1 Flash introduces a new Causal Encoder-Decoder (CED) architecture. The model has 40 Transformer layers organized as a 20-layer causal encoder followed by a 20-layer decoder. The decoder's global KV cache is projected from the final encoder hidden states rather than computed independently at each layer.
This design has a practical consequence: the model activates 8B parameters per token during prefill (input processing) and 16B per token during decode (output generation). For input-heavy workloads — long document analysis, retrieval-augmented generation, agentic tool-calling loops — you pay for far fewer active parameters on the input side than the 552B total might suggest.
The MoE layer uses 1 shared expert and 384 routed experts, activating 6 routed experts per token. The model also includes a 196B-parameter Engram conditional memory, sparsely accessed via token-based lookup.
KV cache compression. V4.1 Flash uses Compressed Sparse Attention 2 (CSA2), which assigns each attention layer one of three static modes — Full, Reindex, or Reuse — to share KV states across layers. Combined with FP4 main KV caching (E2M1 format), the global KV cache footprint drops to about 890 bytes per token. Compared with V4 Flash, that is roughly 1/4 the HBM and 1/8 the SSD storage for the persistent cache.
For API consumers, the KV cache improvement translates directly into lower cache-hit pricing and faster time-to-first-token on long contexts.
Native Multimodal: Vision Input Support
V4.1 Flash processes images natively. Unlike the earlier V4-Flash-Vision-Exp (which bolted a vision encoder onto V4 Flash), V4.1 Flash was trained from scratch as a multimodal model — the vision encoder (DeepSeek-ViT with 2D-RoPE and 3x3 pixel-unshuffle downsampling) and a two-layer MLP projector are part of the original pre-training on 45T tokens.
To send an image:
response = client.chat.completions.create(
model="deepseek-flash",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "Describe this architecture diagram."},
{
"type": "image_url",
"image_url": {"url": "https://example.com/diagram.png"},
},
],
}
],
)
Base64-encoded images are also supported. The vision capability is available on the same deepseek-flash model slug — no separate model name needed.
On DeepSeek's internal benchmarks, V4.1 Flash scores 56.5 on MMMU-Pro, 77.9 on CVBench, and 95.6 on DocVQA (LLM-Judge). V4 Pro does not support vision at all.
Pricing: Peak, Off-Peak, and Cache Economics
V4.1 Flash uses time-of-day pricing. Off-peak rates are 50% of peak rates. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday; all other hours are off-peak.
| Off-Peak | Peak | |
|---|---|---|
| Input (cache miss) | $0.15 / 1M tokens | $0.30 / 1M tokens |
| Input (cache hit) | $0.003 / 1M tokens | $0.006 / 1M tokens |
| Output | $0.60 / 1M tokens | $1.20 / 1M tokens |
Cache-hit input pricing is 98% cheaper than cache-miss input pricing. For agentic workloads that re-send long system prompts or tool definitions across turns, the cache hit rate directly determines your effective cost. The KV cache compression in V4.1 Flash means DeepSeek can offer these low cache-hit rates while still fitting more concurrent contexts per GPU.
Comparison with V4 Pro:
| V4.1 Flash (off-peak) | V4 Pro (off-peak) | |
|---|---|---|
| Input (cache miss) | $0.15 | $0.66 |
| Input (cache hit) | $0.003 | $0.022 |
| Output | $0.60 | $1.98 |
V4.1 Flash is roughly 4x cheaper on input and 3x cheaper on output compared to V4 Pro at off-peak rates — and it outperforms V4 Pro on most benchmarks.
Source: DeepSeek Models & Pricing, retrieved September 11, 2026.
Thinking Mode and Reasoning Capabilities
V4.1 Flash supports both thinking and non-thinking modes. Thinking mode is enabled by default — the model generates chain-of-thought reasoning in a reasoning_content field before producing the final answer.
To disable thinking mode:
response = client.chat.completions.create(
model="deepseek-flash",
messages=[{"role": "user", "content": "Summarize this text."}],
extra_body={"thinking": {"type": "disabled"}},
)
V4.1 Flash also supports continuously controllable reasoning effort from 1 to 100. At maximum effort (100), it achieves a Codeforces rating of 3471 (compared to V4 Pro's 3348), 90.9 on GPQA Diamond, and 65.6 on MathArena Apex.
On agentic benchmarks at maximum effort, V4.1 Flash stands out:
- Terminal-Bench 2.1: 90.6 (ahead of Opus 5.0 at 89.1 and GPT-5.6 Sol at 88.8)
- DeepSWE v1.1: 74.2 (ahead of GPT-5.6 Sol at 73.0)
- AutomationBench: 54.8 (ahead of Opus 5.0 at 50.3)
- Agent's Last Exam: 31.8 (ahead of all listed competitors)
These are vendor-reported benchmarks. Independent evaluations are still emerging as the model just launched.
Source: DeepSeek V4.1 Flash Technical Report, retrieved September 11, 2026.
Supported Features
| Feature | Support |
|---|---|
| JSON Output | Yes |
| Tool Calls / Function Calling | Yes |
| Responses API | Yes |
| Anthropic API format | Yes |
| Chat Prefix Completion (Beta) | Yes |
| FIM Completion (Beta) | Non-thinking mode only |
| Vision | Yes |
| Context Length | 1M tokens |
| Max Output | 384K tokens |
| Concurrency Limit | 2,500 |
DeepSeek also supports the Anthropic messages API format at https://api.deepseek.com/anthropic. If your stack uses the Anthropic SDK, you can point it at DeepSeek without switching to the OpenAI format.
Source: DeepSeek API Docs, retrieved September 11, 2026.
Model Name Migration: What Got Retired
V4.1 Flash replaces two models immediately:
| Retired Model | Legacy Slug | Routing Behavior |
|---|---|---|
| V4 Flash | deepseek-v4-flash | Now routes to V4.1 Flash at V4.1 Flash pricing |
| V4-Flash-Vision-Exp | deepseek-v4-flash-vision-exp | Now routes to V4.1 Flash at V4.1 Flash pricing |
Starting September 14, 2026 at 04:00 UTC, deepseek-v4-pro requests will also route to V4.1 Flash at V4.1 Flash rates. DeepSeek has since clarified that V4 Pro API service will continue after September 14 with unchanged billing, pending further notice.
If your code uses deepseek-v4-flash or deepseek-v4-flash-vision-exp, it will continue to work — the slugs silently resolve to V4.1 Flash. But we recommend updating to deepseek-flash to avoid confusion when V4.1 Pro eventually launches.
Source: DeepSeek V4.1 Flash Announcement, retrieved September 11, 2026.
Common Errors and Gotchas
429 Too Many Requests. V4.1 Flash has a 2,500 concurrent request limit per account. If you hit 429 errors, you can request a capacity expansion at no extra cost. Requests count as one concurrent connection from send to response completion.
Thinking mode on by default. If you are migrating from V4 Flash and expect non-thinking behavior, you must explicitly disable thinking. The default changed — thinking mode is now on for all DeepSeek models.
Legacy model names still work but bill differently. deepseek-v4-flash and deepseek-v4-flash-vision-exp now bill at V4.1 Flash rates, which are lower than V4 Flash's original pricing. Your costs may decrease without code changes.
FIM Completion requires non-thinking mode. Fill-in-the-Middle completion only works when thinking is disabled. Sending a FIM request with thinking enabled returns an error.
Vision input token counting. Images consume tokens based on resolution after downsampling. A high-resolution diagram may use significantly more input tokens than expected. Monitor your token usage when first adopting vision features.
Rate Limits and Concurrency
| Model | Concurrency Limit |
|---|---|
deepseek-flash (V4.1 Flash) | 2,500 |
deepseek-v4-pro | 500 |
Concurrency limits apply at the account level regardless of which API key is used. For per-user isolation, pass a user_id parameter — DeepSeek uses it for content safety isolation, KV cache isolation, and scheduling isolation.
Source: DeepSeek Rate Limit & Isolation, retrieved September 11, 2026.
TheRouter Integration: Routing V4.1 Flash in Fallback Chains
TheRouter routes OpenAI-compatible requests through configured providers, so you can add V4.1 Flash to your routing configuration alongside other models.
A practical pattern: use V4.1 Flash as your default model for general tasks, with a fallback to a different provider if DeepSeek returns a 5xx or times out.
# Example: V4.1 Flash primary, Qwen3.7-Max fallback
models:
- provider: deepseek
model: deepseek-flash
priority: 1
- provider: dashscope
model: qwen3.7-max
priority: 2
V4.1 Flash's 2,500 concurrency limit and low pricing make it a strong candidate for the primary slot in high-throughput routing configurations. For tasks that require vision input, V4.1 Flash is currently one of the cheapest multimodal options available via an OpenAI-compatible API.
For more on fallback routing, see our model fallbacks guide.
Production Checklist
Before deploying V4.1 Flash in production, verify these items:
- Update model slugs. Switch from
deepseek-v4-flash/deepseek-v4-flash-vision-exptodeepseek-flash. - Set thinking mode explicitly. Do not rely on the default. If you need non-thinking behavior, disable it in every request.
- Test vision inputs if migrating from Vision-Exp. V4.1 Flash's vision encoder differs from the experimental version. Run your image-processing test suite against the new model.
- Monitor cache hit rates. V4.1 Flash's compressed KV cache changes cache behavior. Track
prompt_tokens_details.cached_tokensin responses to understand your effective input cost. - Review concurrency headroom. The limit jumped from whatever your V4 Flash allocation was to 2,500 by default. If you were previously hitting limits, you may have room to increase throughput.
- Schedule flexible workloads off-peak. Off-peak pricing is half of peak — batch jobs, evaluations, and non-latency-sensitive tasks should target off-peak hours (all hours outside 01:00–04:00 and 06:00–10:00 UTC on weekdays).
Summary
DeepSeek V4.1 Flash packs a lot into one model: 552B MoE parameters with 8B/16B asymmetric activation, native vision, 1M context, thinking mode with controllable effort, and pricing that starts at $0.15/M input tokens off-peak. It replaces V4 Flash, V4-Flash-Vision-Exp, and (partially) V4 Pro in a single update.
For routing operators, the key facts are: it is OpenAI-compatible at https://api.deepseek.com, the model slug is deepseek-flash, it has a 2,500 concurrency limit, and its agentic benchmark performance rivals frontier models at a fraction of the cost.