Wan 3.0 on VisionX.
One take up to thirty seconds, built from stills, footage and sound at once.
What Wan 3.0 is on the roster
Wan 3.0 is the generation where Alibaba stopped splitting the job across models. Text, a first frame, a first and last frame, reference images, reference clips, reference audio — all of it goes into a single request, and the model works out what kind of shot you are asking for from what you handed it.
That matters most on long output. A take can run to thirty seconds with sound generated alongside the picture, so a complete beat — a line delivered, a move landed, a product turned — comes out whole rather than assembled from three renders in an editor.
Two grades run on the roster. Base 3.0 is the standing choice; 3.0 Prime is the accelerated build of the same model, for when a board is waiting on the take rather than on the budget. Same inputs, same limits, same Cast — the composer quotes both before you run either.
Thirty seconds, one continuous take
Double Wan 2.7’s ceiling, with audio rendered in the same pass. Long enough to carry a beat end to end instead of cutting around the seams.
Ten stills, five clips, five tracks
Every reference pool is separate, so a character, a product, a location, a camera move and a piece of score can all brief the same shot. Wan 2.7 caps images and video at five between them.
One model, every task type
First frame, first-and-last frame, reference-driven, or text alone — there is no mode to pick and no second model id to route to. The references you attach are the instruction.
A 480p rung for probing
The cheapest tier on the Alibaba roster, and the only Wan rung below 720p. Useful for finding the blocking before a take is worth rendering at delivery size.
Choose speed or cost per shot
Prime is the accelerated grade of the same model. Nothing about the shot changes — same references, same ceilings, same output — so the choice is purely how fast you need it back.
Your prompt, unedited
Wan offers to rewrite prompts before rendering. VisionX keeps that off, so what the composer wrote is what runs — no vendor rewrite quietly contradicting your Cast or style.
What Wan 3.0 actually does
What the 3.0 generation does that Wan 2.7 cannot.
A thirty-second take, in one piece
Double 2.7’s ceiling. A complete beat — a line delivered, a move landed, a product turned — renders whole instead of being assembled from three clips in an editor.
Long takeThree reference pools, counted separately
Ten reference images, five reference clips and five audio tracks in one call. Wan 2.7 allows five images and videos between them, so a character, a product, a location and a camera move can finally brief the same shot.
Multimodal referenceNo mode to choose
One model handles text-only, a first frame, a first and last frame, and reference-driven shots alike. It reads which one you meant from the references you attached, so there is no wrong route to pick.
One model idSound rendered with the picture
Audio comes out of the same pass as the frames rather than being dubbed on afterwards, and it costs the same whether you keep it or switch it off.
Joint audio-videoA 480p rung for probing
The only Wan tier below 720p, and the cheapest second of video on the Alibaba roster. Find the blocking there before spending delivery-size renders on it.
Draft tierFraming it picks for itself
Aspect defaults to adaptive, so the model composes for the shot unless you pin it to one of the five fixed ratios. An attached first frame overrides both — the anchor sets the frame.
Adaptive aspectEvery variant, exactly as the studio runs it
These rows are read live from the engine registry and the pricing engine. What you see here is what the composer quotes.
Self-serve rates run as low as $0.085 per VX at scale. The composer quotes the exact VX for your shot, references included, before you run it.
How to write for Wan 3.0
Wan reads references by type and number, so the prompt’s job is to say what each one is for. Everything else is ordinary shot description.
Image 1 is the presenter. Image 2 is the handset. Image 3 is the room. Video 1 carries the camera move — a slow push, then a lift. Audio 1 sets the pace; cut nothing against it. The presenter from Image 1 turns the handset from Image 2 in the room from Image 3, speaking to camera. The move follows Video 1, settling on the handset face-on as the track resolves.
Address references by type and number
“Image 1”, “Video 1”, “Audio 1” — in the order you attached them. Naming what each contributes beats “in the style of”, the same rule the Seedance tiers follow.
A first frame overrules the ratio
Anchor a shot and the output inherits that image’s framing. Compose the anchor at the ratio you want delivered rather than setting one that will be ignored.
Describe the sound you want
Audio renders with the picture, so dialogue, ambience and effects belong in the prompt. Silence is a choice you make by switching audio off, not by omitting it.
Budget the clock when a clip rides along
Reference footage and output share a thirty-second budget, so a shot built on a source clip delivers shorter than one built from stills. Ask for the length you need at the length you can get.
What teams route to Wan 3.0
The board lines that get routed to Wan 3.0 first.
A beat that cannot be cut
Thirty unbroken seconds where a cut would break the illusion — a single move through a space, a monologue, a demo that has to be seen to be believed.
A shot with four separate briefs
This face, that product, this location, that camera move — each from its own reference, none of them averaged into the others.
Cheap blocking before an expensive take
Run the 480p rung until the staging is right, then re-run the same prompt at 1080p.
A deadline rather than a budget
Route to 3.0 Prime when the review is waiting on the take. Same shot, same references, quoted before you run it.
Wan 3.0 against the alternatives
Where 3.0 sits against the engine a Wan board line would otherwise go to, and against the other long-take route on the roster.
| Capability | Wan 3.0 | Wan 2.7 | Seedance 2.5 |
|---|---|---|---|
| Max single take | 30s | 15s | 30s |
| Reference images per call | 10 | 5 (shared with video) | 30 |
| Reference video / audio | 5 clips + 5 tracks | Shared pool of 5, one driving track | Yes |
| Lowest rung | 480p | 720p | 480p |
| Max resolution | 1080p | 1080p | 1080p |
| Audio generation | Same pass, switchable | Same pass | Same pass as the picture |
Every row is read from the VisionX engine registry, which tracks what the composer can actually send — not a vendor deck. Alibaba also documents document and webpage inputs for 3.0; VisionX does not run them, so they are not claimed here. As of August 2026.
Same Cast, same wallet, different look
Route a board line to Wan 3.0, then run the identical shot on another family and compare takes. Identity lives in your Cast, not in any one model, so switching engines never strands the campaign.