How to Prompt Seedance 2.5
A prompt that works fine for a 5-second clip tends to fall apart at 30. The character drifts, the camera wanders, the lighting shifts halfway through for no clear reason. That's not the model losing the plot, it's usually the prompt never having a plot to begin with. ByteDance publishes an actual formula for this, and it's worth following instead of guessing.
The six-part formula
A Seedance 2.5 prompt has six slots, always in this order: Subject, Action or event, Scene and environment, Visual style, Camera movement or cut, Audio.
Only the first two are required. Subject and action are the floor, everything after that is optional detail you add when it actually matters. This is where most prompts go wrong in one of two directions: too thin, just a subject and a vague verb, which gets you something generic, or too stuffed, every cinematography term you know piled into one paragraph, which fights itself. Telling the model "handheld documentary" and "locked-off symmetrical composition" in the same prompt doesn't average out to something interesting, it just gets you mush.
Naming your references
The biggest practical change from Seedance 2.0 isn't the reference count, it's that references are now labeled. Tag each one with @image, @video, or @audio and tell the model what job it does: @image1 defines the subject's face and outfit. @image2 defines the location. Upload eight unlabeled images and the model has to guess which one is the character and which is just style, exactly the failure mode 2.0 was prone to.
For stability, keep it modest even though the ceiling is high: 1 to 8 distinct image subjects, 1 to 5 video subjects at roughly 5 to 10 seconds each, and for editing work, a source clip under 20 seconds with 1 to 5 supporting references. The model can technically take more, it's just less predictable once you do. None of this is really the lever worth pulling, though. A single well-specified reference paired with precise, sensory prompt writing consistently outperforms a stack of references described vaguely.
Staging a full 30 seconds
Don't write one continuous description and expect it to hold for the whole clip. Break the 30 seconds into timestamped beats instead. A simple structure that works for most narrative shots:
0 to 6s: opener, wide shot, establish the scene
6 to 14s: development, medium shot, the main action
14 to 24s: escalation, a moving shot or a cut to a detail
24 to 30s: resolution, close-up, the button on the scene
You don't need to timestamp every second, just the beats that actually matter. Three or four per 15 seconds is a reasonable starting density for straightforward narrative work.
Build in a buffer on any beat involving fast or multi-step action. The model tends to render physical action a little slower than real-world pacing, so a beat you'd expect to take 2 seconds in reality often needs closer to 4 in the prompt.
Directing the audio
Audio generates in the same pass as the video, automatically. Four bracket types keep it clean once a scene has more than one layer of sound: round brackets for music, angle brackets for sound effects, curly braces for spoken dialogue, and full-width square brackets for on-screen subtitles. Plain language works too, this only starts mattering once a prompt has enough going on that things could get crossed, a line of dialogue rendered as an on-screen caption, say. You can also just say what you don't want: "no captions" and "no background music, keep the room sound" are both documented and both work.
Writing like a director
The prompts that actually hold together for the full 30 seconds read like a shot brief, not a description, and none of what makes that true requires more references. A tightly specified single-subject prompt with real sensory detail, exact color, exact motion, beats a five-reference prompt with vague description every time.
Open with a style lock. One line before anything else: genre, color grade, film stock or digital look, and what should never appear. Everything after it has to match, and this single line does more for consistency across 30 seconds than almost anything else you can add.
Tie the camera to an event, not a clock. "Push in the moment she looks up" holds together better than "push in from 10 to 15 seconds," because the model is following an action rather than a countdown. Save exact timestamps for the handoffs that genuinely need to land on a beat, an entrance, an exit, a hard cut.
Lock what shouldn't move. A short line at the end naming the one or two things that have to stay exactly as they are, a jacket color, the framing, the light, is often the difference between a clip that holds together and one that quietly drifts by the last third.
Editing a clip you already have
Seedance 2.5 can revise part of a finished clip instead of regenerating the whole thing, worth knowing before you reflexively re-roll a near-miss. Feed the clip in as @video1, name exactly what changes, and lock everything that shouldn't move. Editing keeps the original aspect ratio and roughly the original duration.
Edit @video1. Keep the subject, background, camera movement, and overall color grading exactly as they are. Replace only the wristwatch's face between 4 and 7 seconds, restoring a clean metallic dial with no flicker or distortion at the edges of the edit.
That's the whole grammar: name the target, say what it becomes, lock the rest.
Five prompts that work
Five complete prompts across different jobs, each built around one style lock, camera tied to an event, and an explicit continuity line, not a longer reference list. Swap the subject and references for your own and run as-is.
Neon train dance
Text-to-video, 10 seconds
Style: High-energy futuristic music video, neon rim lighting, anamorphic lens flare, glossy chrome surfaces Scene: The interior of a high-speed maglev train car at night, city lights streaking past the windows in motion blur Subject: Two dancers in reflective silver outfits, defined by @image1 and @image2 Action: They perform sharp, synchronized choreography, hitting a hard pose exactly on each bass drop Camera: Static wide shot for the first half, whip-pan into a low-angle close-up on the final pose Audio: (a driving electronic bass track), a sharp impact sound synced to each pose hit, no dialogue Continuity: Keep both dancers' outfits and the train's neon lighting identical throughout, no color drift
Rooftop to rooftop
Image-to-video, 8 seconds
Style: High-speed action, motion blur, desaturated steel-blue palette, documentary-style camera shake Scene: The roof of a maglev train tearing through a mountain corridor, defined by @image1 for the character Subject: A figure in a weatherproof jacket, running along the train's roof Action: They sprint the length of the roof, leap across the gap between cars, and land in a roll as the train enters a tunnel Camera: Fast tracking shot alongside the run, whip-pan to follow the jump, holding steady through the landing Audio: (wind roar and train engine noise), a sharp impact on landing, no music Continuity: Keep the jacket color and the train's roof geometry consistent throughout
The rune awakening
Image-to-video, 10 seconds
Style: Dark fantasy, cold rim lighting, slate and ash palette with faint cyan emissive accents Scene: A stone chamber wall etched with runes, defined by @image1 for the character Subject: A robed sorceress, downcast gaze, cold light on her face Action: She slowly raises her head to meet the camera as the runes behind her ignite one by one Camera: Slow dolly back as her gaze lifts, ending on a three-quarter portrait Audio: (a low resonant hum building as each rune ignites), breath visible in the cold air, no dialogue Continuity: Keep the rune pattern and her exact facial features consistent as the light builds
The bodega run
Image-to-video, 30 seconds
Style: 30-second continuous single take, natural real-time speed, no cuts Scene: a small New York City bodega on a rainy morning, defined by @image1 Subject: a bike messenger in a yellow rain jacket, carrying one red bicycle helmet in the left hand throughout, defined by @image2 Action: 0-5s: the door bell rings as the messenger enters, closes the glass door with the right hand, and shakes rain from the shoulders without dropping the helmet. 5-10s: the camera follows from behind at chest height as the messenger walks to the drink cooler, opens it with the right hand, removes one clear bottle of seltzer, and closes the door. 10-16s: the messenger turns toward the counter and walks around one stationary customer without changing hands, red helmet still in the left, bottle still in the right. 16-22s: sets the bottle on the counter, taps a phone once on the card reader, waits for one confirmation beep, then picks the bottle back up. 22-27s: the clerk gives a small nod, the messenger turns back toward the entrance as the camera backs up to a medium full-body frame. 27-30s: opens the door with the right forearm, exits into the rain, the door closes behind, camera holds inside the store. Camera: continuous handheld tracking shot from behind at chest height for the full 30 seconds, no cuts Audio: (rain outside, door bell, cooler hum, footsteps, bottle set on counter, one card-reader beep, quiet store room tone), no music Continuity: keep the same messenger, jacket, helmet, bottle, phone, clerk, counter, cooler, and store layout from first frame to last. No repeated entrance, no duplicated bottle or helmet, no object teleportation.
The bottle's four worlds
Image-to-video, 30 seconds
Style: premium brand film, cinematic hard cuts on the beat, rich color grading that shifts per world, glossy production value Scene: begins in a minimalist white product studio Subject: a glowing glass orb, defined by @image1, that travels through a sequence of transforming environments Action: 0-5s: the orb rests on a white pedestal in the studio, light catching its surface as the camera holds steady. 5-11s: the studio dissolves into a moonlit forest, the orb now resting on a moss-covered stone, fireflies drifting past it. 11-17s: the forest gives way to a rain-soaked city street at night, the orb held in an open palm, neon reflections crossing its surface. 17-23s: the scene shifts to a warm desert dune at golden hour, the orb balanced on sand, heat shimmer visible behind it. 23-30s: the orb returns to the original white studio pedestal, camera pulling back to a wide final shot. Camera: slow continuous push-in through the first four scenes, pulling back to a wide shot only in the final beat Audio: (a building orchestral score that shifts instrumentation with each world), no dialogue, no captions Continuity: keep the orb's exact size, glow color, and surface detail identical across every environment, only the surroundings change
Common questions
Do I have to fill in all six parts of the prompt?
No. Subject and Action are the only required parts. Add Scene, Style, Camera, and Audio only when you need to control them.
What happens if I write one long paragraph instead of staging it with timestamps?
It usually still generates, but a 30-second stream-of-consciousness prompt tends to produce a 30-second stream-of-consciousness result. Staging with timestamps gives the model clearer instructions for how the scene should progress.
Does the reference-tagging syntax work on Seedance 2.0 too?
Yes, the @image, @video, and @audio tagging carries over from 2.0. What's different on 2.5 is how many references you can attach at once. hat's different on 2.5 is how many references you can attach at once, though as above, that's rarely the thing worth maximizing.
Do I need a lot of references to get a good result?
No, and this is the most common misunderstanding. One well-specified reference with exact color, motion, and style detail consistently outperforms several references described vaguely. Reference count isn't the lever that matters.
Can I fix one part of a clip without regenerating the whole thing
Yes. Feed the finished clip in as @video1, name exactly what changes, and lock everything that shouldn't move. It keeps the original aspect ratio and roughly the original duration.



