ByteDance Flagship

Seedance 2.0

ByteDance's next-generation AI video model with the revolutionary @-reference system. Combine text, images, video clips, and audio in a single prompt. Native audio-video synchronization, V2V editing, and up to 2K resolution at 30fps — all in one unified generation.

About

About Seedance 2.0

Seedance 2.0 is ByteDance's most advanced AI video generation model, unveiled in February 2026. It adopts a unified multimodal audio-video joint generation architecture supporting 4 input modalities simultaneously — text, up to 9 images, up to 3 video clips, and up to 3 audio tracks. The ground-breaking @-reference system lets you tag specific elements in your prompt and bind them to uploaded references for granular control over camera movement, character appearance, audio rhythm, and visual style. Outputs reach up to 2K resolution with native synchronized audio including multilingual lip-sync, sound effects, and background music.

About Seedance 2.0

Key Features of Seedance 2.0

@-reference system, 4-modality input, native audio sync, V2V editing, 2K resolution

Core Features Overview

@-Reference System

The @-reference system is Seedance 2.0's signature innovation. You tag elements in your text prompt with @Image1, @Video1, @Audio1 and bind them to uploaded reference files. The model extracts camera movements from video references, beat rhythms and voice characteristics from audio references, and composition/styles from image references. This enables unprecedented control: 'A @Image1 walks through @Image2 while @Video1 camera movement plays and @Audio1 sets the mood.'

Prompt
Output (Example)

@Image1 walks through @Image2 with camera movement from @Video1 and background music from @Audio1

Multi-reference prompt combining all modalities

@-Reference System Example

@Image1 character dances with rhythm from @Audio1 in @Image3 environment

Character motion guided by audio beat reference

@-Reference System Example

Native Audio Generation

Seedance 2.0 uses a dual-branch diffusion transformer architecture that processes video and audio latents in parallel with shared cross-attention. This means audio is not added as a post-processing step — it is generated simultaneously with the visuals, ensuring millisecond-level synchronization. The model can generate multilingual lip-sync dialogue, action-matched sound effects, and mood-appropriate background music, all controlled through text prompts or audio references.

Prompt
Output (Example)

A person giving a presentation with synchronized English speech and slide transitions

Lip-sync dialogue with visual content

Native Audio Generation Example

Cooking tutorial with step-by-step narration and ambient kitchen sounds

Narration synchronized with cooking actions

Native Audio Generation Example
Official Showcase

Official Showcase

Explore Seedance 2.0's capabilities in multimodal reference control, native audio generation, and video editing

Multi-reference prompt combining all modalities@-Reference System

@Image1 walks through @Image2 with camera movement from @Video1 and background music from @Audio1

Multi-reference prompt combining all modalities

Character motion guided by audio beat reference@-Reference System

@Image1 character dances with rhythm from @Audio1 in @Image3 environment

Character motion guided by audio beat reference

Lip-sync dialogue with visual contentNative Audio Generation

A person giving a presentation with synchronized English speech and slide transitions

Lip-sync dialogue with visual content

Narration synchronized with cooking actionsNative Audio Generation

Cooking tutorial with step-by-step narration and ambient kitchen sounds

Narration synchronized with cooking actions

FAQ

Seedance 2.0 FAQ

Seedance 2.0 FAQ

The @-reference system lets you tag elements in your prompt with @Image1, @Video1, @Audio1 labels and bind them to uploaded reference files. Seedance 2.0 extracts camera movements from video references, beat rhythms from audio, and composition styles from images. This gives you granular control over every aspect of the generated video.

Seedance 2.0 supports 4 input modalities simultaneously: text prompts (unlimited length), up to 9 reference images (≤30MB each), up to 3 video clips (2-15s total duration, ≤50MB each), and up to 3 audio tracks (≤15s total, ≤15MB each). Total file limit: 12 files per request.

Seedance 2.0 outputs at native 2K (2048x1080) resolution at 30fps with multiple quality levels: 480p, 720p, and 1080p. Video duration ranges from 4 to 15 seconds per generation. Supported aspect ratios include landscape, portrait, and 21:9 ultra-wide.

Seedance 2.0 uses a dual-branch architecture that processes video and audio latents in parallel. Audio is generated simultaneously with visuals, ensuring millisecond-level synchronization. It supports multilingual lip-sync dialogue, action-matched sound effects, and mood-appropriate background music. You can also upload audio references as input.

V2V editing allows you to upload existing video clips as reference and generate new videos that inherit their motion patterns, camera paths, and pacing. You can change specific elements like outfits, actions, or scene details while preserving the original motion structure.

Seedance 2.0 adds video and audio reference inputs, increases image references from 1 to 9, introduces the @-reference system for multimodal control, adds V2V video editing, extends max resolution from 1080p to 2K, increases duration from 12s to 15s, and is approximately 30% faster than 1.5 Pro.

Seedance 2.0 uses per-second dynamic pricing based on resolution: 480p (14-28 credits/second), 720p (28.5-57 credits/second), and 1080p (640-3,810 credits/second). There are two speed variants: Standard and Fast, with Fast being roughly 30% faster.

Seedance 2.0 is ideal for video directors needing precise motion control, content creators wanting native audio sync without post-production, advertisers producing branded video content, educators creating narrated tutorials, and anyone who needs professional-quality AI video with synchronized sound.

Testimonials

What Creators Say About Seedance 2.0

The @-reference system is genuinely revolutionary. I can extract camera movements from a reference clip and apply them instantly — it's a completely new creative workflow.

Alex Kim

Alex Kim

Video Director

Alex Kim: “The @-reference system is genuinely revolutionary. I can extract camera movements from a reference clip and apply them instantly — it's a completely new creative workflow.

Priya Sharma: “Native audio sync saves hours of post-production. The lip-sync quality is surprisingly precise even with non-English dialogue.

Lucas Müller: “V2V editing lets me enhance existing footage without reshooting. Seedance 2.0 is now a core tool in our production pipeline.

Yuki Tanaka: “The 4-modality input is a game-changer. I can bring a character design, a camera movement reference, and background music all into one prompt and get exactly what I envisioned.

Explore More AI Video Models

Start Creating with Seedance 2.0

Experience Seedance 2.0 — the most advanced video generator from ByteDance, free online

user 1
user 2
user 3
user 4
user 5

10,000+ users