ByteDance · BytePlusVIDEOCore tiersNative audio

Seedance 2.0 on VisionX.

Four generations of one model family, from prompt probe to native 4K master.

4models on rosterfrom 5 VXper shotup to 9reference imagesCast-readyidentity locked

CINEMA TIER · RENDERED ON VISIONX

01

What Seedance 2.0 is on the roster

THE MODEL

Seedance is the spine of VisionX video. Four generations of the same model family are live at once, so a shot moves from a cheap probe to a mastered final without changing platforms, prompts, or Cast.

Seedance 1.0 lite and 1.5 pro exist to be spent freely: test a prompt, find the blocking, throw the takes away. Seedance 2.0 fast is the production default. Seedance 2.0 is the mastering engine, and it renders native 4K directly, with no upscale pass.

One family, draft to master

Look-dev on the cheap engines carries straight up the ladder. The prompt that worked on Seedance 1.5 pro works on 2.0.

The deepest reference support on the roster

Seedance 2.0 takes up to 9 reference images and accepts video input, which is what Cast identity and Reference Plates run on.

Native multi-clip extend

Seedance 2.0 and 2.0 fast fuse up to 3 finished clips into one continuous video, so a 5-second hero moment becomes a scene.

02

What Seedance 2.0 actually does

CAPABILITIES

Six things the Seedance tiers do that a text-prompt-only model cannot.

Twelve references, and each one has a job

Nine reference images plus video and audio slots feed a single Seedance 2.0 generation. Each asset carries a role — this face, that product, this camera move — instead of being blended into an average.

Multimodal reference

Sound generated with the picture, not over it

Dialogue, ambience and effects come out of the same pass as the frames, so lip movement is shaped by the waveform rather than dubbed onto a finished clip afterwards.

Joint audio-video

Several cuts inside one render

A single generation can hold more than one shot, keeping one character, one location and one soundtrack across the cuts rather than stitching three clips together in an editor.

Multi-shot

Copy the camera move from a clip you upload

Feed a reference video and the model reads its move — the push, the orbit, the whip — then applies that motion to your subject instead of to your footage.

Video-to-video

Lip-synced dialogue across languages

Write the line in quotes and the character speaks it. English and Mandarin hold up best; short lines map to phonemes most cleanly.

Multilingual speech

Native 4K on the mastering tier

Seedance 2.0 renders 4K directly out of the model, so a hero shot never goes through an upscale pass that softens what the grade is meant to hold.

Delivery resolution
03

Every variant, exactly as the studio runs it

SPEC SHEET

These rows are read live from the engine registry and the pricing engine. What you see here is what the composer quotes.

ModelModesResolutionsDurationsVX
Seedance 1.0 liteSeedance 1.0 lite tierT2V · I2V480p · 720p5s / 10s6 VX · 5s T2V at 720p
Seedance 1.5 proSeedance 1.5 pro tierT2V · I2V480p · 720p · 1080p5s / 10s5 VX · 5s T2V at 720p
Seedance 2.0 fastSeedance 2.0 fast tierT2V · I2V · V2V480p · 720p5s / 10s / 15s8.5 VX · 5s T2V at 720p
Seedance 2.0Seedance 2.0 tierT2V · I2V · V2V480p · 720p · 1080p · 4K5s / 10s / 15s11 VX · 5s T2V at 720p

Self-serve rates run as low as $0.085 per VX at scale. The composer quotes the exact VX for your shot, references included, before you run it.

04

How to write for Seedance 2.0

PROMPTING

Seedance addresses uploads by tag. Naming the role is the difference between a reference the model uses and one it averages away.

PRODUCT AD · MULTI-SHOT · 16:9
Use [Image1] as the bottle. Keep the label sharp and unchanged.
Use [Image2] for the set: wet black stone, one hard key light from camera left.
Match the camera move in [Video1] — slow push, slight handheld drift.
Take the room tone and pace from [Audio1].

Shot 1 — macro on condensation.
Shot 2 — pull back, bottle enters frame.
Shot 3 — hold, label centred.

Give every slot a role, not a mood

"Use [Image1] as the character's face" beats "in the style of [Image1]". The model resolves roles precisely and moods loosely.

Number your shots inside one prompt

Multi-shot is generated jointly, so listing the cuts gets you a sequence that holds continuity — not three clips you have to match afterwards.

Put spoken lines in quotes

Quoted text is read as dialogue and drives lip-sync. Keep it short and name the language for the cleanest result.

Look-dev cheap, master late

Find the blocking on Seedance 1.0 lite or 1.5 pro, then send the prompt that worked up to 2.0 fast or 2.0. The ladder is the same model family, so the take carries.

05

What teams route to Seedance 2.0

USE CASES

The board lines that get routed to Seedance first.

Scroll-stopping social hooks

Vertical shots with a spoken opening line and ambience already in the mix — ready for Reels, Shorts and TikTok without a separate audio pass.

Product ads from a single packshot

Drop the pack into an image slot and a camera-move clip into a video slot. The label stays legible through the whole move.

Previz and storyboard motion

Turn boards into moving shots that hold the same character across cuts, so a director can read blocking before anything is booked.

Talking-head presenters

One front-facing portrait plus a script line gives a presenter whose mouth matches audio the model generated itself.

Music-led visuals

Load a track into an audio slot and the cuts and motion take their timing from it rather than from a fixed cadence.

One ad, several markets

Re-run the same shot with the line rewritten per language. Lip movement regenerates with the speech instead of drifting against a dub.

06

Seedance 2.0 against the alternatives

COMPARISON

The short version against the two engines it gets compared with most. All three run on the same VisionX board, so you can settle it on your own footage.

Seedance 2.0 capability comparison against Kling 3.0 and Veo 3.1
CapabilitySeedance 2.0Kling 3.0Veo 3.1
Reference input typesText, image, video, audioText, imageText, image
Audio generationSame pass as the pictureSeparate sync stepSame pass as the picture
Audio file as an inputYesNoNo
Multi-shot in one generationYesPartialPartial
Native 4K without an upscaleYes, on Seedance 2.0Vendor-announcedNo

VisionX rows are read from the engine registry. Competitor rows describe the vendors’ own published capabilities and are current as of July 2026 — vendors ship fast, so re-check before quoting these in a deck.

FAQ

Seedance 2.0, answered

How many references can Seedance take in one generation?
Up to twelve alongside your text prompt on Seedance 2.0: nine reference images, plus video clips or audio files sharing the remaining slots. You address them in the prompt by tag, such as [Image1] or [Video1], and each one carries the role you name rather than being averaged into the others.
Does Seedance generate audio and lip-sync?
It generates sound in the same pass as the picture on Seedance 2.0 and 2.0 fast, so dialogue, ambience and effects are timed to the frames rather than added afterwards. Lip-synced speech works across several languages, with English and Mandarin the most reliable. Put the spoken line in quotation marks in your prompt.
What is the difference between the four Seedance tiers?
They are four generations of the same model family, at four grades of spend. Seedance 1.0 lite and 1.5 pro are cheap enough to throw away — use them to find the prompt and the blocking. Seedance 2.0 fast is the production default. Seedance 2.0 is the mastering engine and renders native 4K directly, with no upscale pass. A prompt that works on a cheap engine carries up the ladder.
Can Seedance make a video longer than one generation?
Yes. Seedance 2.0 and 2.0 fast fuse finished clips into one continuous video natively, so a hero moment can be extended into a scene rather than re-generated at a longer duration and re-graded from scratch.
Can I feed Seedance an existing video clip?
Yes — Seedance 2.0 and 2.0 fast accept video input. The common use is handing the model a clip whose camera move you want, so it reads the push, orbit or whip and applies that motion to your subject instead of to the uploaded footage.
Should I use Seedance 2.0 or Seedance 2.5?
Both are live, so route by the shot. Seedance 2.5 is the long-form engine: one continuous take up to 30 seconds, 30 reference images per call, and prompt-driven editing and extension — at up to 1080p. The 2.0 engines remain the draft-to-master ladder, and Seedance 2.0 is still the only Seedance engine that renders native 4K, so hero masters stay there. Because identity lives in your Cast rather than in a model, moving a shot between them is a routing choice, not a rebuild. The Seedance 2.5 page covers the long-take engine in detail.

Try Seedance 2.0 on your Cast.

20 VX free on signup, no card required.