AlibabaVIDEOGuest engineNative audio

Wan 3.0 on VisionX.

One take up to thirty seconds, built from stills, footage and sound at once.

2models on rosterfrom 8.75 VXper shotup to 10reference imagesCast-readyidentity locked
01

What Wan 3.0 is on the roster

THE MODEL

Wan 3.0 is the generation where Alibaba stopped splitting the job across models. Text, a first frame, a first and last frame, reference images, reference clips, reference audio — all of it goes into a single request, and the model works out what kind of shot you are asking for from what you handed it.

That matters most on long output. A take can run to thirty seconds with sound generated alongside the picture, so a complete beat — a line delivered, a move landed, a product turned — comes out whole rather than assembled from three renders in an editor.

Two grades run on the roster. Base 3.0 is the standing choice; 3.0 Prime is the accelerated build of the same model, for when a board is waiting on the take rather than on the budget. Same inputs, same limits, same Cast — the composer quotes both before you run either.

Thirty seconds, one continuous take

Double Wan 2.7’s ceiling, with audio rendered in the same pass. Long enough to carry a beat end to end instead of cutting around the seams.

Ten stills, five clips, five tracks

Every reference pool is separate, so a character, a product, a location, a camera move and a piece of score can all brief the same shot. Wan 2.7 caps images and video at five between them.

One model, every task type

First frame, first-and-last frame, reference-driven, or text alone — there is no mode to pick and no second model id to route to. The references you attach are the instruction.

A 480p rung for probing

The cheapest tier on the Alibaba roster, and the only Wan rung below 720p. Useful for finding the blocking before a take is worth rendering at delivery size.

Choose speed or cost per shot

Prime is the accelerated grade of the same model. Nothing about the shot changes — same references, same ceilings, same output — so the choice is purely how fast you need it back.

Your prompt, unedited

Wan offers to rewrite prompts before rendering. VisionX keeps that off, so what the composer wrote is what runs — no vendor rewrite quietly contradicting your Cast or style.

02

What Wan 3.0 actually does

CAPABILITIES

What the 3.0 generation does that Wan 2.7 cannot.

A thirty-second take, in one piece

Double 2.7’s ceiling. A complete beat — a line delivered, a move landed, a product turned — renders whole instead of being assembled from three clips in an editor.

Long take

Three reference pools, counted separately

Ten reference images, five reference clips and five audio tracks in one call. Wan 2.7 allows five images and videos between them, so a character, a product, a location and a camera move can finally brief the same shot.

Multimodal reference

No mode to choose

One model handles text-only, a first frame, a first and last frame, and reference-driven shots alike. It reads which one you meant from the references you attached, so there is no wrong route to pick.

One model id

Sound rendered with the picture

Audio comes out of the same pass as the frames rather than being dubbed on afterwards, and it costs the same whether you keep it or switch it off.

Joint audio-video

A 480p rung for probing

The only Wan tier below 720p, and the cheapest second of video on the Alibaba roster. Find the blocking there before spending delivery-size renders on it.

Draft tier

Framing it picks for itself

Aspect defaults to adaptive, so the model composes for the shot unless you pin it to one of the five fixed ratios. An attached first frame overrides both — the anchor sets the frame.

Adaptive aspect
03

Every variant, exactly as the studio runs it

SPEC SHEET

These rows are read live from the engine registry and the pricing engine. What you see here is what the composer quotes.

ModelModesResolutionsDurationsVX
Wan 3.0Wan 3.0T2V · I2V · V2V480p · 720p · 1080p5s / 10s / 15s / 30s8.75 VX · 5s T2V at 720p
Wan 3.0 PrimeWan 3.0 PrimeT2V · I2V · V2V480p · 720p · 1080p5s / 10s / 15s / 30s12.25 VX · 5s T2V at 720p

Self-serve rates run as low as $0.085 per VX at scale. The composer quotes the exact VX for your shot, references included, before you run it.

04

How to write for Wan 3.0

PROMPTING

Wan reads references by type and number, so the prompt’s job is to say what each one is for. Everything else is ordinary shot description.

PRODUCT FILM · ONE TAKE · 30S · ADAPTIVE
Image 1 is the presenter. Image 2 is the handset. Image 3 is the room.
Video 1 carries the camera move — a slow push, then a lift.
Audio 1 sets the pace; cut nothing against it.

The presenter from Image 1 turns the handset from Image 2 in the room
from Image 3, speaking to camera. The move follows Video 1, settling
on the handset face-on as the track resolves.

Address references by type and number

“Image 1”, “Video 1”, “Audio 1” — in the order you attached them. Naming what each contributes beats “in the style of”, the same rule the Seedance tiers follow.

A first frame overrules the ratio

Anchor a shot and the output inherits that image’s framing. Compose the anchor at the ratio you want delivered rather than setting one that will be ignored.

Describe the sound you want

Audio renders with the picture, so dialogue, ambience and effects belong in the prompt. Silence is a choice you make by switching audio off, not by omitting it.

Budget the clock when a clip rides along

Reference footage and output share a thirty-second budget, so a shot built on a source clip delivers shorter than one built from stills. Ask for the length you need at the length you can get.

05

What teams route to Wan 3.0

USE CASES

The board lines that get routed to Wan 3.0 first.

A beat that cannot be cut

Thirty unbroken seconds where a cut would break the illusion — a single move through a space, a monologue, a demo that has to be seen to be believed.

A shot with four separate briefs

This face, that product, this location, that camera move — each from its own reference, none of them averaged into the others.

Cheap blocking before an expensive take

Run the 480p rung until the staging is right, then re-run the same prompt at 1080p.

A deadline rather than a budget

Route to 3.0 Prime when the review is waiting on the take. Same shot, same references, quoted before you run it.

06

Wan 3.0 against the alternatives

COMPARISON

Where 3.0 sits against the engine a Wan board line would otherwise go to, and against the other long-take route on the roster.

Wan 3.0 capability comparison against Wan 2.7 and Seedance 2.5
CapabilityWan 3.0Wan 2.7Seedance 2.5
Max single take30s15s30s
Reference images per call105 (shared with video)30
Reference video / audio5 clips + 5 tracksShared pool of 5, one driving trackYes
Lowest rung480p720p480p
Max resolution1080p1080p1080p
Audio generationSame pass, switchableSame passSame pass as the picture

Every row is read from the VisionX engine registry, which tracks what the composer can actually send — not a vendor deck. Alibaba also documents document and webpage inputs for 3.0; VisionX does not run them, so they are not claimed here. As of August 2026.

FAQ

Wan 3.0, answered

What is the difference between Wan 3.0 and Wan 3.0 Prime?
Speed and price. Prime is Alibaba’s accelerated build of the same model — identical inputs, identical reference limits, identical thirty-second ceiling and identical 1080p top end. It costs more per second, and the composer quotes both before you run either, so the choice is simply whether the take or the budget is the constraint.
How is Wan 3.0 different from Wan 2.7?
Three things. The take runs to thirty seconds instead of fifteen. The reference pools are separate and much deeper — ten images, five clips and five audio tracks, against 2.7’s five images and videos combined. And there is no mode to route to: 2.7 splits the work across three models, while 3.0 works out the task from the references you attach.
Does Wan 3.0 render 4K?
No. It tops out at 1080p, and this page will not claim otherwise. Native 4K on VisionX means Seedance 2.0, which renders it directly out of the model with no upscale pass.
Can I feed Wan 3.0 a video clip?
Yes, up to five reference clips totalling fifteen seconds. Note that reference footage and output share a thirty-second budget, so a shot built on a clip delivers a shorter take than one built from stills — the composer accounts for that before it quotes you.
Does VisionX let Wan rewrite my prompt?
No. Wan offers a prompt-rewrite step before rendering and VisionX keeps it switched off, on 3.0 exactly as on 2.7. What the composer wrote is what runs — a vendor rewrite is how a Cast or style note gets quietly contradicted.

Try Wan 3.0 on your Cast.

20 VX free on signup, no card required.