DeepSeek V4-Flash-Vision-Exp Adds Vision to the DeepSeek API: What Your Routing Layer Must Handle Now

DeepSeek launched its first multimodal model on August 21. The critical routing problem is not the model itself — it's that vision requests require a content array where your proxy may only be passing a string.

Published via DeepSeek

Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

Abstract technical diagram showing image tokens flowing through an API routing layer with split paths for text and multimodal requests

DeepSeek released deepseek-v4-flash-vision-exp on August 21 — its first model that accepts image inputs alongside text. In benchmark terms, it keeps parity with V4-Flash on pure-text agent tasks and brings visual agent benchmarks close to Opus-4.8. For routing teams, the benchmark row is the least important thing in the release. The routing problem comes earlier, at the message shape.

What changed at the API level

The model name is deepseek-v4-flash-vision-exp. It lives at the same https://api.deepseek.com base URL, with the same OpenAI-compatible /chat/completions interface and the same Anthropic-compatible /anthropic/messages interface.

The difference is the content field. Text-only requests typically pass content as a string:

{"role": "user", "content": "Summarize this document."}

Vision requests must pass content as an array of typed blocks:

{
  "role": "user",
  "content": [
    {"type": "text", "text": "What does this chart show?"},
    {"type": "image_url", "image_url": {"url": "https://example.com/chart.png"}}
  ]
}

Any proxy middleware that normalizes content to a string — or that wraps only the string case — will send a malformed request and get a 400 back. That includes any thin OpenAI-proxy that was written before multimodal was common, and any homegrown gateway code that does content = str(message) for logging or token estimation.

The Anthropic-compatible endpoint uses a different block shape: instead of image_url, it uses an image block with a source object typed as base64, url, or file. If you're using DeepSeek's Anthropic-format endpoint for routing alongside Claude, you need to handle two distinct image shapes in the same pipeline.

Three image delivery modes, each with different routing implications

DeepSeek supports three ways to send an image, and the right choice for your routing layer is not obvious from the docs alone.

Inline base64 embeds the image data in the request body. It counts toward the 48 MiB request body limit, and each image may be at most 32 MiB. This is the lowest-latency option for single requests — no pre-upload step — but it blows up request sizes and will serialize poorly across a rate-limited proxy.

External URL lets the DeepSeek model download the image itself. The URL must be under 8,192 characters and the image under 32 MiB. The model has 60 seconds to retrieve it. If your images live behind auth or a CDN that rate-limits crawlers, requests will fail intermittently. This is convenient for public assets but unreliable for anything behind a signed URL with a short TTL.

Files API file_id is the right choice when you're sending the same image across multiple requests, or when the image exceeds 32 MiB. Files uploaded via the Files API can be up to 64 MiB and are not counted against the inline body limit. The tradeoff: you need an upload step before the first request, and file IDs are tied to your DeepSeek account — you cannot pass them through a generic proxy that substitutes credentials.

The maximum images per request is 600, with a total request size ceiling of 200 MiB when using file IDs (64 MiB otherwise). Each image is billed as tokens: DeepSeek normalizes every image to roughly 800×800 before counting, so a 5,000×5,000 image costs the same as a 1,000×1,000 one. The per-image token ceiling is 384.

The detail parameter is a cheap triage lever

For image_url inputs, DeepSeek exposes a detail field with values low, high, original, and auto. Setting detail: "low" downscales the image to 512×512 before inference, making it faster and cheaper when pixel-level precision is not required.

This matters for routing. If your pipeline inspects uploaded documents for type-detection — "is this a form, a chart, or a photograph?" — before deciding how to process it, you can route the triage call with detail: "low" and the full-resolution call only when the triage result warrants it. The cost difference can be substantial for high-volume pipelines where most uploads are simple forms or screenshots.

detail: "high" and detail: "original" are equivalent: both keep the original image. detail: "auto" currently resolves to original.

What the experimental status means for your routing policy

The exp suffix is load-bearing. This model does not carry a production SLA. DeepSeek has not published a rate limit tier for it. That means you should not route it as a drop-in replacement for V4-Flash without a fallback.

A workable policy: treat deepseek-v4-flash-vision-exp as the primary route for vision-required requests, with a fallback to another vision-capable provider (Gemini 3.5 Flash, Vertex Imagen endpoint, or any model with image_url support) when you receive a 429, a 503, or a model-not-found error. Keep V4-Flash as the text-only route for the same prompt, so that if vision is optional for your use case, you can degrade gracefully to the text path without an error.

Migration shape: before and after

Before (text-only routing, works with any model):

response = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=[{"role": "user", "content": "Summarize this."}]
)

After (vision routing, model must be the vision variant):

response = client.chat.completions.create(
    model="deepseek-v4-flash-vision-exp",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "Summarize this document."},
            {"type": "image_url", "image_url": {"url": image_url, "detail": "low"}}
        ]
    }]
)

Images are supported only in user messages. System and assistant messages with images return 400. If your agent uses a system-prompt image (for example, a branded header or a reference diagram), that will need to move to the first user turn.

What TheRouter users should watch

Vision routing is still early for DeepSeek. Operators currently have to pick the model name explicitly — there is no automatic routing based on whether a message contains images. A reasonable pattern right now: inspect the message at the gateway layer before dispatch, check whether any content block has type: "image_url" or type: "image", and route to deepseek-v4-flash-vision-exp (or your preferred vision provider) accordingly.

The experimental status and unknown rate limit tier argue for keeping a multi-provider fallback in place rather than treating this as a stable primary route. See the DeepSeek Vision guide for the full limits table.

Help & contact