Back to Models

Doubao Seed 1.6 Vision

doubaodoubao/doubao-seed-1-6-vision

Doubao Seed 1.6 Vision β€” vision-language Seed 1.6 variant for multimodal understanding.

Doubao Seed 1.6 Vision (doubao-seed-1.6-vision-250815) is ByteDance's general-purpose multimodal model exposed via Volcano Engine. It belongs to the Seed1.6 series, a sparse MoE family (230B total / 23B active parameters) that incorporates multimodal understanding during pre-training and supports a 256K context window upstream. The Vision variant adds native VisualCoT β€” the model can crop, rotate, zoom, and annotate images as part of its reasoning chain.

For TheRouter operators, Doubao Seed 1.6 Vision is a cost-effective Chinese-provider vision model with strong GUI-agent, grounding, and video-understanding capabilities. It is fully OpenAI-compatible through api.therouter.ai/v1 and supports temperature, max_tokens, top_p, tools/tool_choice, response_format, and stop β€” exactly the parameters listed in standard-models.yaml.

Best for
  • β€’ GUI agents and grounding tasks where the model must interact with screenshots, PDFs, or web UIs and produce explicit visual reasoning traces.
  • β€’ Video understanding and long-context multimodal workflows (education, image moderation, inspection, AI search) where Chinese-language context is prevalent.
Reach for something else if
  • β€’ Open-weight self-hosting β€” Seed1.6 Vision is API-only; use open VLMs (Qwen3-VL, InternVL) when weights are required.
  • β€’ Pure English multimodal workloads where Gemini 2.5 Pro or Claude Sonnet 4.6 deliver higher measured English VLM benchmarks.
Context Length
131K
Max Output
33K
Input Priceper 1M tokens
$0.1296/ 1M tokens
Output Priceper 1M tokens
$1.30/ 1M tokens

Modalities

textimage→text

Capabilities

Vision

Pricing Breakdown

TypeRate
Input$0.1296 / 1M tokens
Output$1.30 / 1M tokens

Supported Parameters

temperaturemax_tokenstop_ptoolstool_choiceresponse_formatstop

Specifications

Release dateAugust 2025 (250815 snapshot)ohmygpt.com β†—verified
Architecture230B total / 23B active MoE with multimodal pre-trainingseed.bytedance.com β†—verified
VisualCoTNative tool-based visual chain-of-thought (crop, rotate, zoom, annotate)developer.volcengine.com β†—verified
LicenseProprietary (Volcano Engine API)www.volcengine.com β†—verified
Training cutoffNot publicly disclosedunknown
Inference backendsVolcano Engine (proprietary)www.volcengine.com β†—verified

Benchmarks

BenchmarkDistributionScoreSource
MME-SCI (image-only)
Science-focused multimodal benchmark; closed-source Doubao-Seed-1.6 outperformed most open VLMs but trailed top Gemini-class models.
41.32%%arxiv.org β†—
Gaokao 2025 (multimodal)
Seed1.6-Thinking ranked 2nd in science and 1st in humanities among five leading reasoning models on Shandong Gaokao 2025 papers.
676 (science) / 683 (humanities)seed.bytedance.com β†—

API Usage Examples

Use the global api.therouter.ai endpoint shown below for new integrations; the legacy China accelerated endpoint is retired.

cURL
curl https://api.therouter.ai/v1/chat/completions   -H "Content-Type: application/json"   -H "Authorization: Bearer $THE_ROUTER_API_KEY"   -d '{
    "model": "doubao/doubao-seed-1-6-vision",
    "messages": [
      {"role": "user", "content": "Summarize the key points from this input."}
    ]
  }'

Chat completion

cURL
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"doubao/doubao-seed-1-6-vision","messages":[{"role":"user","content":[{"type":"text","text":"描述这张图片"},{"type":"image_url","image_url":{"url":"https://therouter.ai/assets/vision-sample.png"}}]}]}'

More from doubao

Similar models

Cross-provider sibling models

News & changes

2025-06-11

ByteDance announces Seed1.6 series at FORCE conference

ByteDance Seed team unveiled the Seed1.6 family with 256K context, multimodal pre-training, AdaCoT adaptive thinking, and strong Gaokao/JEE results. The Vision variant followed in subsequent months.

re-authored by TheRouterseed.bytedance.com β†—
2025-08

Doubao-Seed-1.6-vision publicly documented (250815)

Volcano Engine documentation and developer evaluations highlight native VisualCoT, 256K context, and strong performance on GUI grounding and video understanding scenarios.

re-authored by TheRoutervolcengine.com β†—

Frequently asked

Does Doubao Seed 1.6 Vision support the same context window as the upstream 256K model?

The upstream Volcano Engine endpoint supports 256K. TheRouter exposes the model with the context_length configured in standard-models.yaml (131072 tokens). Use the YAML value for production capacity planning.

re-authored by TheRouterwww.volcengine.com β†—
How does VisualCoT differ from standard vision tool calling?

VisualCoT is native: the model itself decides when and how to crop/rotate/annotate the image inside its reasoning trace and returns the transformed image as part of the visible CoT. No separate vision-tool microservice is required.

re-authored by TheRouterdeveloper.volcengine.com β†—
Fact ledger β€” every claim on this page traces here
sourceURLretrieved
Release dateohmygpt.com β†—2026-05-25verified
Architectureseed.bytedance.com β†—2026-05-25verified
VisualCoTdeveloper.volcengine.com β†—2026-05-25verified
Licensewww.volcengine.com β†—2026-05-25verified
Training cutoffβ€”β€”unknown
Inference backendswww.volcengine.com β†—2026-05-25verified
MME-SCI (image-only)arxiv.org β†—2026-05-25to verify
Gaokao 2025 (multimodal)seed.bytedance.com β†—2026-05-25to verify
ByteDance announces Seed1.6 series at FORCE conferenceseed.bytedance.com β†—2026-05-25verified
Doubao-Seed-1.6-vision publicly documented (250815)volcengine.com β†—2026-05-25verified
Does Doubao Seed 1.6 Vision support the same context window as the upstream 256K model?www.volcengine.com β†—2026-05-25to verify
How does VisualCoT differ from standard vision tool calling?developer.volcengine.com β†—2026-05-25to verify
Help & contact