Back to Models

Doubao Seedance 2.0

doubaodoubao/doubao-seedance-2-0

Doubao SeedDance 2.0 β€” text/image-to-video generation, flagship quality tier.

Doubao Seedance 2.0 is ByteDance's flagship multimodal video generation model, developed by Team Seed and released through the Volcengine Ark platform. Unlike text-to-video models that only accept text prompts, Seedance 2.0 supports simultaneous text, image, audio, and video inputs β€” up to 12 reference files per request β€” via a Dual-Branch Diffusion Transformer architecture that generates synchronized video and audio in parallel.

The model excels at complex motion generation with realistic physics, native audio-video synchronization (dialogue, sound effects, ambient sound, and music), first-and-last-frame bookend control, and Director Mode for camera trajectory and lighting control via natural language. Output reaches up to 15 seconds at 2K resolution with 24 fps and aspect ratios from 1:1 to 21:9. The API API follows an async submit-poll-download pattern typical for video generation.

Best for
  • β€’ Short-form video content β€” social media clips, advertisements, and promotional videos with native audio
  • β€’ Storyboard prototyping and pre-visualisation β€” animating from first-and-last-frame bookends with cinematic camera motion
  • β€’ Multimodal character-driven videos β€” combining character portraits, motion reference clips, and environment images in one request
  • β€’ Lip-synced talking heads in 8+ languages β€” dialogue generation with native audio alignment
Reach for something else if
  • β€’ Long-form narrative video (over 15 seconds) β€” chain multiple Seedance clips or route to a specialised pipeline
  • β€’ Real-time interactive video generation β€” Seedance has 30–120 second generation latency; use a real-time diffusion model instead
  • β€’ Very high-volume, low-cost thumbnail or short GIF generation β€” dedicated image models (e.g. doubao/doubao-seedream-4-5) are cheaper per output
Context Length
--
Max Output
--
Video Priceper second of video
$0.162/ second

Modalities

textimage→video

Capabilities

VisionText to VideoImage to Video

Media Generation Capabilities

video_generation
sizes
  • 1280x720
  • 720x1280
durations_seconds
  • 5
  • 10
defaults
size
1280x720
duration_seconds
5

Pricing Breakdown

TypeRate
Video$0.162 / second

request is price per second for 720x1280 or 1280x720 outputs β€” the vendor's own per-second rate for this catalog configuration

Supported Parameters

promptimage_urlsecondssizeaspect_ratio

Specifications

Release date2026-02-10 (consumer); 2026-04-02 (API GA via Volcengine Ark)apidog.com β†—verified
ArchitectureDual-Branch Diffusion Transformer (audio + video joint generation)seed.bytedance.com β†—to verify
Max duration15 secondsapidog.com β†—verified
Max resolution2K (2560Γ—1440)www.nxcode.io β†—verified
Frame rate24 fpsapidog.com β†—verified
Input modalitiesText, image (up to 9), video (up to 3 clips), audio (up to 3 files)apidog.com β†—verified
Output audioNative β€” dialogue, sound effects, ambient sound, music; synchronous with videoseed.bytedance.com β†—verified
Language supportMultilingual (Chinese + English native), lip-sync in 8+ languagesapidog.com β†—verified
Model versiondoubao-seedance-2-0-260128apidog.com β†—verified
LicenseCommercial use permitted via Volcengine Ark / TheRouter APIvolcengine.com β†—to verify

Benchmarks

BenchmarkDistributionScoreSource
SeedVideoBench-2.0 (internal)
ByteDance's internal SeedVideoBench-2.0 evaluates motion stability, prompt adherence, temporal consistency, and audio-video synchronisation. Seedance 2.0 reports state-of-the-art results. No third-party audit available.
SOTA (internal benchmark)seed.bytedance.com β†—
Vbench
Third-party reviews (e.g., evoLink, WaveSpeed) rank Seedance 2.0 highly on motion naturalness and audio sync, comparable to Kling 3.0 and Sora 2 in overall quality.
4.2+/5.0 (third-party estimate)evolink.ai β†—

API Usage Examples

Use the global api.therouter.ai endpoint shown below for new integrations; the legacy China accelerated endpoint is retired.

cURL
curl https://api.therouter.ai/v1/videos   -H "Content-Type: application/json"   -H "Authorization: Bearer $THE_ROUTER_API_KEY"   -d '{
    "model": "doubao/doubao-seedance-2-0",
    "prompt": "A slow tracking shot through a neon-lit city street",
    "duration": 8
  }'

API guide

Text-to-video generation

Submit a text prompt to generate a video with native audio. The API uses an async task pattern β€” submit, poll for completion, then download the result URL.

cURL
# Step 1: Submit text-to-video generation task
curl -X POST https://api.therouter.ai/v1/contents/generations/tasks \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "doubao/doubao-seedance-2-0",
    "content": [
      {
        "type": "text",
        "text": "A golden retriever running through a sunlit wheat field, wide tracking shot, cinematic lighting"
      }
    ],
    "resolution": "1080p",
    "ratio": "16:9",
    "duration": 5,
    "watermark": false
  }'

# Step 2: Poll for completion (save task_id from response)
curl https://api.therouter.ai/v1/contents/generations/tasks/cgt-xxxxxx \
  -H "Authorization: Bearer $THEROUTER_API_KEY"

# Step 3: When status="succeeded", download video_url from the response

More from doubao

Similar models

Cross-provider sibling models

News & changes

2026-04-14

Volcengine Ark fully opens Seedance 2.0 API services to enterprise and individual users

Volcengine officially launched Seedance 2.0 API with four input methods (text, image, audio, video), establishing safety standards for portrait/copyright compliance, and deploying over 10,000 pre-installed virtual avatars. Early enterprise adopters report 80–90% efficiency gains in short drama and manga-style video production workflows accessible through TheRouter.

re-authored by TheRouteraibase.com β†—
2026-02-10

ByteDance releases Seedance 2.0 with Director Mode, native audio, and 2K output

Seedance 2.0 introduced a Dual-Branch Diffusion Transformer architecture for joint audio-video generation, Director Mode for precise camera trajectory and lighting control via natural language prompts, first-and-last-frame bookend control, and lip-sync in 8+ languages. The model went viral for its realistic physics and motion stability, quickly becoming one of the most widely discussed AI video models of 2026.

re-authored by TheRouterreddit.com β†—

Frequently asked

What input types does Seedance 2.0 accept?

Seedance 2.0 accepts text, images (up to 9, max 30MB each), video clips (up to 3, 2–15 seconds each, max 50MB each), and audio files (up to 3, MP3, max 15MB each) β€” all in a single request. You can mix and match modalities for multimodal reference, e.g., a character image + motion reference video + ambient audio + text prompt.

re-authored by TheRouterapidog.com β†—
How long does video generation take?

Typically 30–120 seconds depending on the resolution and duration. A 5-second 1080p video averages 60–90 seconds. The API uses an async pattern: submit a task, poll the status endpoint every 5–10 seconds, and download the result URL when status reaches 'succeeded'. The URL expires after 24 hours.

re-authored by TheRouterapidog.com β†—
Can I use Seedance 2.0 with the standard OpenAI SDK?

No β€” Seedance 2.0 uses Volcengine Ark's async task API, not OpenAI's chat completions format. TheRouter routes requests to the correct Ark endpoint, but you must use the async submit-poll-download pattern. Native OpenAI SDK methods like chat.completions.create() will not work. Refer to the API guides above for the correct request format.

re-authored by TheRouterapidog.com β†—
What's the difference between Seedance 2.0 and Seedance 2.0 Fast?

Seedance 2.0 Fast is a lower-cost, faster variant optimised for iterative video drafting and prototyping. It uses a smaller model that generates videos more quickly and at a lower per-request price (3.0556 vs 3.8889). The trade-off is reduced output quality, particularly in motion complexity, prompt adherence, and audio-video synchronisation. For production-grade deliverables, use the standard Seedance 2.0; for quick drafts and internal iterations, Fast is more economical.

re-authored by TheRouterwww.volcengine.com β†—
Fact ledger β€” every claim on this page traces here
sourceURLretrieved
Release dateapidog.com β†—2026-05-26verified
Architectureseed.bytedance.com β†—2026-05-26to verify
Max durationapidog.com β†—2026-05-26verified
Max resolutionwww.nxcode.io β†—2026-05-26verified
Frame rateapidog.com β†—2026-05-26verified
Input modalitiesapidog.com β†—2026-05-26verified
Output audioseed.bytedance.com β†—2026-05-26verified
Language supportapidog.com β†—2026-05-26verified
Model versionapidog.com β†—2026-05-26verified
Licensevolcengine.com β†—2026-05-26to verify
SeedVideoBench-2.0 (internal)seed.bytedance.com β†—2026-05-26single source
Vbenchevolink.ai β†—2026-05-26single source
Volcengine Ark fully opens Seedance 2.0 API services to enterprise and individual usersaibase.com β†—2026-05-26verified
ByteDance releases Seedance 2.0 with Director Mode, native audio, and 2K outputreddit.com β†—2026-05-26verified
What input types does Seedance 2.0 accept?apidog.com β†—2026-05-26to verify
How long does video generation take?apidog.com β†—2026-05-26to verify
Can I use Seedance 2.0 with the standard OpenAI SDK?apidog.com β†—2026-05-26to verify
What's the difference between Seedance 2.0 and Seedance 2.0 Fast?www.volcengine.com β†—2026-05-26to verify
Help & contact