Aristotto
Back
Guides

WAN 3.0: What It Does and How to Prompt It

Aristottoby Aristotto11 min

Wan 3.0 is Alibaba's latest video generation model, officially launched on August 24, 2026 after a public beta from August 6. It generates up to 30 seconds in a single pass with native synchronized audio, accepts text, images, video, audio, and runs on Aristotto with the 46% per-generation credit reduction through December 20, 2026.

This is both a review and a prompting guide, because understanding what this model is good at tells you what to prompt it for, and understanding how it reads a prompt tells you why certain shots fail when they shouldn't.

What is Wan 3.0?

Wan 3.0 is a multimodal video generation model from Alibaba's Tongyi Lab. It takes text prompts, up to 10 reference images, up to 5 reference videos, up to 5 reference audio files, and uniquely, documents. Hand it a PDF, a slide deck, a spreadsheet, or a webpage, and it builds a video sequence from the content inside. No other model does this as of August 2026.

It generates at 480p, 720p, or 1080p across five aspect ratios (16:9, 9:16, 1:1, 4:3, 3:4), from 2 to 30 seconds, with audio composed in the same pass as the video: dialogue with lip-sync across 12 languages, ambient sound, music, and sound effects.

It also includes an AI Director mode for scripting up to 6 distinct shots with individual camera angles and time ranges in a single generation, with transitions handled automatically. And there's a Prime tier for higher fidelity on demanding shots.

No 4K output. No open weights for 3.0.

Where Wan 3.0 is strongest

It protects the source image. When you feed it a reference image and a prompt, it treats the image as ground truth. The character's face, outfit, the room's layout, the product's color hold through motion, camera moves, and the full 30 seconds. Independent same-seed testing against Seedance 2.0 confirms this across dozens of cases. For product video, talking-head content, or any job where the reference is the deliverable and the model's job is to make it move without changing what it looks like, this is the model's best case.

Physics reads as real. Glass shatters along actual fracture lines. Billiard balls transfer momentum believably. Water tracks with gravity. Independent testing gave Wan 3.0 a clear, reproducible edge here over both its predecessor and its closest competitor.

The subject stays locked while the world changes. Ask the sky to shift color while a mountain stays still, and most models let the mountain wobble. Wan 3.0 holds the foreground rock-solid while the background animates. That separation matters for any shot where a product, a person, or a building needs to stay perfectly still while the environment moves around it.

Document-to-video is genuinely new. A quarterly report becomes a 30-second animated summary without anyone describing what the charts look like. A slide deck becomes a narrated walkthrough without rewriting each slide as a prompt.

Where Wan 3.0 still breaks

Uncommanded camera cuts. Ask for one continuous camera move and Wan 3.0 sometimes inserts a hard, unrequested shot change, with a visible color-grade jump or brief ghosting. This is documented across multiple independent tests and is the model's most significant structural flaw right now. Review every clip for uninvited cuts before sending it anywhere.

It flinches at surreal prompts. Ask for indoor rain in a sunny living room and it negotiates the instruction down to a lighting change rather than committing. It sides with the image over the prompt, which is exactly the instinct that makes it so reliable for faithful work, and exactly why it's the wrong model for imagination-driven creative. If the concept requires the image to become something it's not, use a different model.

Multi-subject scenes are weaker. Two characters interacting, passing an object, sharing a space with distinct blocking, is where Wan 3.0 falls behind competitors that handle multi-character work more reliably.

How to structure a Wan 3.0 prompt

Every Wan 3.0 prompt follows a five-segment structure. The order matters because the model applies heavier weight to what appears first and last.

Shot type and scene, then camera movement, then subject and appearance, then audio, then @references.

Lead with the shot type: wide, medium, close-up, extreme close-up, aerial, POV, over-the-shoulder, Dutch angle. This anchors the spatial composition before anything else. Then name the camera move using standard terms: slow push in, tracking shot, orbit, crane up, handheld, static locked. Generic descriptions like "moving camera" produce inconsistent results.

Then the subject with specific physical detail, then the setting and lighting, then audio as its own distinct layer. End with any @reference tags linking uploaded assets (@Image1 through @Image9, @Video1 through @Video3, @Audio1 through @Audio3).

What you leave unnamed, the model invents. Every layer you name is one less thing the model guesses.

Tight beats outperform dense description. 50 to 120 words is the working range. A prompt with 5 to 10 specific things that happen, in order, outperforms one packed with 40 adjectives. Past a certain density, the model's attention narrows to whatever is stated most concretely and repeatedly, and softer clauses drop out. When a result is close but wrong, name the specific miss rather than rewriting everything.

Directing audio in Wan 3.0

Audio generates in the same pass as the video. Name each layer as its own directive, ranked by prominence: foreground dialogue, ambient bed, background detail. Use verbs tied to visible events ("the sharp click of the latch releasing") rather than mood adjectives ("realistic door sounds").

For spoken dialogue: quote the exact line, name a visible speaker on camera with enough detail to pin who's talking, and end with "no subtitles" so the model renders the line as audio rather than text baked into the frame. Lip-sync works at the phoneme level across 12 languages including English, Spanish, French, German, Japanese, Korean, and Mandarin.

Six prompts across different jobs

Each one follows the official five-segment structure. Type and duration are platform settings, not prompt text.

Food, cooking

Text-to-video, 10 seconds

Medium shot of a woman in a clay-stained linen apron standing at a thick wooden farmhouse table, morning light flooding in from a large window behind her, catching flour dust suspended in the air. She lifts a sheet of fresh pasta from the table with both hands, arms extended, the pasta draping in a long translucent sheet, light passing through it and casting a warm amber glow on her face and forearms. Camera holds static at chest height then begins a slow push-in as she folds the sheet carefully onto a floured wooden board, dusting it once with her right hand, flour cloud catching the backlight. Her expression is focused and unhurried, not performing for camera, working. Kitchen behind her is lived-in: copper pots on the wall, a half-empty bottle of olive oil, a torn paper bag of flour. Warm golden natural light, shallow depth of field, background soft. Audio: the soft thud of dough on wood, the whisper of flour scattering, a faint creak of the wooden table, birdsong through the open window

Short film, four-shot narrative

Text-to-video, 28 seconds

Shot 1 [0-7s]: Wide establishing shot of a rain-soaked train station platform at night, empty benches, amber platform lights reflecting off wet concrete, no characters visible yet.
Shot 2 [7-15s]: Medium shot, a man in a dark coat enters from the right, walks slowly to the end of the platform, stops, checks his watch, looks down the empty tracks.
Shot 3 [15-22s]: Close-up on his face as a distant train light appears, his expression shifts from waiting to recognition, rain on his shoulders.
Shot 4 [22-28s]: Wide pullback, the platform from above as the train pulls in, he steps toward the doors, one other passenger steps off.
Overall tone: sparse ambient score, no dialogue, naturalistic sound design, rain throughout, distant train brakes in Shot 4. Muted desaturated color grade, 35mm film grain.

Social content, vertical

Text-to-video, 6 seconds

Close-up of hands peeling a ripe mango over a wooden cutting board in bright natural window light, juice running down the fingers. Camera holds static with a slight handheld sway. Audio: the wet peel of the skin separating, the soft thud of the pit hitting the board, no music, no dialogue. Vivid saturated warm color grade.

Fitness, boxing

Text-to-video, 10 seconds

Medium tracking shot of a boxer in a grey tank top and dark shorts working a heavy leather bag in a well-lit modern gym. Camera tracks alongside at chest height, handheld energy. He throws a fast rhythmic combination of jabs and crosses, sweat visible on his shoulders, the bag swinging between strikes. Audio: the sharp thud of each strike landing on leather, his controlled exhale on every punch, the metallic creak of the bag chain, no music, no dialogue. High-contrast color grade.

Multi-shot product launch

Text-to-video, 24 seconds

Shot 1 [0-6s]: Aerial shot, overhead view of a matte-black wireless speaker on a white marble surface, soft studio light, camera slowly descends.

Shot 2 [6-12s]: Medium shot, hands carefully lift the speaker from minimal packaging, crisp unboxing sounds.

Shot 3 [12-18s]: Close-up, camera orbits the speaker at surface level, macro detail on the fabric mesh and brushed-metal controls.

Shot 4 [18-24s]: Wide shot, speaker placed on a warm wooden shelf in a living room, camera pulls back slowly to reveal the full space, warm evening light. Overall tone: minimal ambient sound design, no dialogue, subtle orchestral swell at Shot 4. Product reference: @Image1.

Anime scene

Reference-to-video, 8 seconds

Medium shot of the character from @Image1, short silver hair, dark teal jacket, square messenger bag slung across one shoulder. She walks from frame left into a quiet train platform, stops beside the yellow safety line, turns toward camera as wind catches her hair. Camera holds static. Clean ink linework, flat cel shading, muted color palette matching the reference sheet, consistent line thickness throughout. Audio: distant platform announcement echo, wind through the station, no music, no dialogue, no subtitles. 1080P.

Anime and stylized content

Wan 3.0 handles anime and illustrated styles, but drift is more visible here than in photorealistic content. An anime face that shifts line thickness or eye shape between frames reads immediately as broken, where a photorealistic face can shift subtly and the viewer might not notice.

Build a locked reference package first: front character sheet, three-quarter pose, face close-up, background plate, and prop sheet. Reuse it across every generation. Change only one variable per test (seed, prompt wording, duration, or audio instruction) so you can trace what caused a failure. Review at three checkpoints per clip: opening frame, midpoint, and final second.

The 30-second window

It works. Identity holds across the full duration. Multi-stage action sequences execute in order. That's real.

The caveat is the uncommanded camera-cut issue. The longer the clip, the more chances the model takes to insert an uninvited shot change. Until that's resolved, review every 30-second clip for surprise cuts. The duration is genuine, but "30 seconds" and "one unbroken shot" are not guaranteed to be the same thing.

Beat count matters more than raw seconds. Roughly 5 to 8 seconds per beat keeps pacing natural. Requesting 30 seconds for a single small moment leaves the model padding time.

What does Wan 3.0 cost on Aristotto?

Wan 3.0 is available on Aristotto with the 46% per-generation credit reduction running on all paid tiers through December 20, 2026. That's the same reduction as Seedance, applied to credit cost per generation, not the plan price.

Common questions

What is Wan 3.0?

Wan 3.0 is Alibaba's multimodal video generation model, launched August 24, 2026. It generates up to 30 seconds of video with native audio from text, images, video, audio, and document inputs.

Is Wan 3.0 available on Aristotto?

Yes, with the 46% per-generation credit reduction on all paid tiers through December 20, 2026.

How long can Wan 3.0 generate in a single clip?

to 30 seconds per generation.

Does Wan 3.0 support 4K output?

No. It generates at 480p, 720p, and 1080p across five aspect ratios.

What is Wan 3.0 AI Director mode?

Wan 3.0 AI Director is an advanced multi-shot sequencing engine that automatically generates up to six continuous cinematic cuts from a single prompt. Unlike basic text-to-video tools, it autonomously structures camera paths, handles transitions, maintains visual continuity, and builds a multi-angle scene over a 30-second timeline.

Is Wan 3.0 AI Director the same as keyframing?

No, Wan 3.0 AI Director and keyframing are completely different video control methods. Keyframing requires a user to upload explicit static images (such as First-Last-Frame interpolation) to establish manual start and endpoints, whereas the AI Director relies on semantic text or scripts to automatically dictate creative multi-shot camera logic and timing.

How should I structure a Wan 3.0 prompt?

Follow the five-segment order: shot type and scene, camera movement, subject and appearance, audio direction, @reference tags. Lead with shot type to anchor composition.

What languages does Wan 3.0 lip-sync support?

Phoneme-level lip-sync across 12 languages including English, Spanish, French, German, Japanese, Korean, and Mandarin.

What is Wan 3.0's biggest weakness?

Uncommanded camera cuts: the model sometimes inserts hard, unrequested shot changes during continuous takes, documented across independent tests.

How does Wan 3.0 compare to Seedance 2.5?

Both generate 30-second clips. Wan 3.0 is stronger on faithful image-to-video and physics. Seedance 2.5 is stronger on creative interpretation and multi-subject scenes. Choose by the shot.

Discover more

View all