Back to Models

Doubao 1.5 Vision Pro 32k

doubaodoubao/doubao-1-5-vision-pro-32k

Doubao 1.5 Vision Pro (32k context) โ€” extended-context vision-language variant.

Doubao-1.5-vision-pro-32k is ByteDance's flagship vision-language model released alongside Doubao 1.5 Pro in January 2025 (model id suffix 250115). It extends the sparse MoE architecture of the 1.5 Pro family with comprehensive upgrades in multimodal data processing, dynamic resolution handling, and fine-grained visual-text alignment, enabling strong performance on text-rich image understanding and visual reasoning tasks.

For TheRouter operators, Doubao-1.5-vision-pro-32k offers a cost-efficient Chinese-first multimodal endpoint that supports standard OpenAI vision chat format. It excels at document QA, chart/table extraction, OCR-heavy workflows, and GUI understanding at dramatically lower price than GPT-4o or Claude 3.5 Sonnet vision variants. The model is fully reachable through api.therouter.ai/v1 with the same production inference stack (PD-disaggregation, W4A8 quantization) used by the text-only 1.5 Pro sibling.

Best for
  • โ€ข Chinese document QA, invoice/contract extraction, chart & table understanding where native Chinese OCR and layout comprehension matter.
  • โ€ข GUI / mobile app understanding and visual agent workflows that require fine-grained screenshot reasoning in Chinese UI contexts.
Reach for something else if
  • โ€ข Pure English high-resolution image captioning or artistic image generation where Gemini-2.5 or GPT-4o vision variants currently lead measured benchmarks.
Context Length
33K
Max Output
16K
Input Priceper 1M tokens
$0.486/ 1M tokens
Output Priceper 1M tokens
$1.46/ 1M tokens

Modalities

textimageโ†’text

Capabilities

Vision

Pricing Breakdown

TypeRate
Input$0.486 / 1M tokens
Output$1.46 / 1M tokens

Supported Parameters

temperaturemax_tokenstop_ptoolstool_choiceresponse_formatstop

Specifications

Release date2025-01-15www.aibase.com โ†—verified
ArchitectureSparse MoE vision-language model (extension of Doubao-1.5-pro)team.doubao.com โ†—verified
Input modalitiesText + Image (no native video)www.volcengine.com โ†—verified
LicenseClosed (API only)verified
Training cutoffNot publicly disclosedunknown

Benchmarks

BenchmarkDistributionScoreSource
DocVQA
Strong text-rich document understanding; one of the highest reported scores among 2025 MLLMs.
96.7%%arxiv.org โ†—
InfoVQA
Excellent infographic and chart reasoning performance.
89.3%%arxiv.org โ†—
VisuLogic
Highest score among compared models on this visual reasoning benchmark (still 23+ points below human).
28.1%%arxiv.org โ†—
ChartQA
Competitive chart and table extraction; exact numeric scores reported in OCR-Reasoning paper.
Strongarxiv.org โ†—

API Usage Examples

Use the global api.therouter.ai endpoint shown below for new integrations; the legacy China accelerated endpoint is retired.

cURL
curl https://api.therouter.ai/v1/chat/completions   -H "Content-Type: application/json"   -H "Authorization: Bearer $THE_ROUTER_API_KEY"   -d '{
    "model": "doubao/doubao-1-5-vision-pro-32k",
    "messages": [
      {"role": "user", "content": "Summarize the key points from this input."}
    ]
  }'

API guide

Vision chat completion

cURL
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model":"doubao/doubao-1-5-vision-pro-32k",
    "messages":[{"role":"user","content":[
      {"type":"text","text":"่ฏท็”จไธญๆ–‡ๆ่ฟฐ่ฟ™ๅผ ๅ›พ็‰‡ไธญ็š„่กจๆ ผๅ†…ๅฎนใ€‚"},
      {"type":"image_url","image_url":{"url":"https://therouter.ai/assets/vision-sample.png"}}
    ]}]
  }'

More from doubao

Similar models

Cross-provider sibling models

News & changes

2025-01-15

ByteDance releases Doubao-1.5-vision-pro alongside Doubao 1.5 Pro

ByteDance launched the Doubao visual understanding model Doubao-1.5-vision-pro with comprehensive upgrades in multimodal data processing and dynamic resolution. The model targets text-rich image reasoning and Chinese document workflows at the same aggressive pricing as the text-only 1.5 Pro sibling.

re-authored by TheRouteraibase.com โ†—

Frequently asked

Does Doubao-1.5-vision-pro-32k support the standard OpenAI image_url format?

Yes. It accepts the same chat.completions payload structure with image_url content blocks as other OpenAI-compatible vision models exposed through TheRouter.

Fact ledger โ€” every claim on this page traces here
sourceURLretrieved
Release datewww.aibase.com โ†—2026-05-26verified
Architectureteam.doubao.com โ†—2026-05-26verified
Input modalitieswww.volcengine.com โ†—2026-05-26verified
Licenseโ€”โ€”verified
Training cutoffโ€”โ€”unknown
DocVQAarxiv.org โ†—2026-05-26verified
InfoVQAarxiv.org โ†—2026-05-26verified
VisuLogicarxiv.org โ†—2026-05-26verified
ChartQAarxiv.org โ†—2026-05-26to verify
ByteDance releases Doubao-1.5-vision-pro alongside Doubao 1.5 Proaibase.com โ†—2026-05-26verified
Does Doubao-1.5-vision-pro-32k support the standard OpenAI image_url format?src โ†—2026-05-26to verify
Help & contact