Aristotto
Back
Guides

Flux 3: What It Does and How to Prompt It

Aristottoby Aristotto12 min

Flux 3 is Black Forest Labs' video generation model, launched in early access on August 4, 2026. It does something architecturally different from every other video model covered on this blog: it trains image, video, and audio on one shared set of weights rather than generating them separately and syncing them afterward. That single fact changes what it's good at and how you should prompt it.

It generates clips from 5 to 20 seconds with native audio produced in the same pass as the video at no extra cost. Multilingual dialogue with phoneme-level lip-sync is supported natively. Available on Aristotto with 46% off, through December 20, 2026.

What is Flux 3?

Flux 3 is a multimodal generation model from Black Forest Labs. Where most video models generate footage and then add audio through a separate pass or pipeline, Flux 3 learns all three modalities, still image, motion, and sound, as one shared representation. The practical result: when a foot hits the ground, the footstep sound lands on that exact frame. When glass shatters, the crack aligns with the fracture. When a door closes, the click matches the hand. The model learned these as the same event, not two things to synchronize after the fact.

It supports five workflows: text-to-video, image-to-video, video-to-video, keyframes (up to 10 ordered stills with frame positions that the model interpolates into continuous motion), and video continuation (extending an existing clip).

One more thing worth knowing: Flux 3 will default to producing multiple scenes and camera angles within a single generation unless you explicitly ask for one continuous shot. That's a feature, not a bug, but it surprises people coming from models that only produce one locked-off take per generation.

Where Flux 3 is strongest

Audio that's physically correct, not just present. This is the headline capability. Impacts, footsteps, engines, weather, breaking objects, fabric tearing, anything with a visible cause and an audible effect comes back with the sound precisely matched to the visual event. Other models approximate this through a sync pass. Flux 3 produces it natively.

Style range. This is where Flux 3 genuinely pulls ahead of most competitors. It doesn't default to one polished cinematic look. Ask for claymation and you get visible fingerprint textures on the clay, tool marks, wobbly stop-motion steps. Ask for VHS found footage and you get autofocus hunting, interlaced scan lines, blown highlights from porch lights, the way a cheap camcorder actually behaves. Ask for comic-book halftone and you get ink outlines, dot shading, chromatic aberration splits. Most video models can attempt stylized work. Flux 3 commits to it.

Multi-scene direction in one pass. Where most models produce one continuous shot, Flux 3 is built to handle cuts, angle changes, and scene transitions within a single generation. Write HARD CUT between timestamped ranges and the model executes distinct camera setups in sequence. A 10-second clip can carry an establishing wide, a cut to a close-up, and a pull-back to a reveal without stitching.

Keyframe interpolation. Feed it up to 10 ordered stills with specific frame positions, and it generates the continuous motion between them. You specify every composition. The model fills in only the movement. No other model on this list offers this as a distinct workflow.

Transformations. Pin a start image and an end image, and the model fills the middle with a continuous physical transformation. A rusted car restores itself panel by panel. A man in a coat becomes something else entirely. The two stills carry most of the outcome, meaning the quality depends more on how well the frames match than on how you word the prompt.

Where Flux 3 still falls short

20 seconds is the ceiling. Seedance 2.5 and Wan 3.0 both generate 30. That's enough for most jobs, but it's not the full 30-second continuous take that some workflows now depend on.

No reference tagging. Seedance and Wan let you upload multiple references and tag each one by role. Flux 3 uses one image as the first frame, or up to 10 ordered keyframes, but there's no equivalent of the tagged multi-reference system. Omni-reference is announced but not live.

No targeted editing. If a clip is 90% right, you re-generate or extend. You can't fix one specific moment the way Seedance 2.5's timestamp editing allows. Video editing is listed as coming soon.

Early access depth. The community testing base is much thinner than what Seedance or Kling have built over six months of production use.

How Flux 3 reads a prompt

Flux 3 prompts are written as natural prose, not labeled fields. Black Forest Labs' own guidance: direct the scene, don't inventory the objects in it.

Describe what happens, how the subject moves, what the camera does, and what it sounds like. Concrete nouns and verbs that a camera could actually resolve. Adjectives alone leave the result to chance. And crucially, describe the audio too, not just the picture. Telling the model what things sound like ("the shriek of tearing metal," "the low buzz of the neon transformer") changes the output, not just the soundtrack.

Four prompt shapes, each for a different job:

A short phrase works when there's one clear subject and you want the model to fill in everything else. Fast for exploration, risky when a specific detail has to survive.

A natural-language paragraph is the default. Camera, then subject, then action, then environment, then sound. One flowing passage.

Labeled fields break the same content into named lines. Worth the extra structure when you're tuning one variable at a time.

Timestep prompting splits the runtime into ranges with one beat each, and marks angle changes with HARD CUT on its own line. This is the shape for anything that has to land on a specific mark. Selecting a frame range also sets how many seconds each beat gets, so the timing and the composition are one decision.

Dialogue goes in double quotes. Close with "no on-screen text, no subtitles" so the model renders spoken lines as audio rather than text baked into the frame. Multilingual dialogue is supported natively, including lip-sync across languages.

Over-stuffing hurts. Black Forest Labs is explicit that too much detail can make motion less coherent. The workflow you choose already answers some questions, so spend the words on what it doesn't cover.

Directing audio in Flux 3

Audio generates automatically in every clip at no extra cost.

Because the model learns sound and motion together, the most effective audio direction is physical: describe the events that produce sound rather than the sound itself. "The crunch of gravel under boots" gives the model a visible action to sync to. "Atmospheric ambient sound" gives it nothing specific.

Scenes with clear physical events produce the strongest audio. If a result comes back thin on sound, add a physical event the model can see and match, rather than an adjective describing a mood.

For complex sound design, rank the layers by prominence: foreground action first, ambient bed second, background detail third.

Four prompts that work

The second moon, Text-to-video, 10 seconds

A family camping trip beside a lake at night, shot on a cheap camcorder in 1995. The camera starts pointed at the campfire, then slowly pans to the black water. A second moon rises out of the lake and hangs just above the far shore, its reflection stretching toward the camera. Everyone at the camp runs toward the water and the camera shakes violently, the flashlight blows out the foreground. The moon folds back into the lake and vanishes. One continuous take, soft autofocus hunting, interlaced analog noise, muffled voices and splashing footsteps. No music, no edits, no cinematic polish.

The meadow clouds, Text-to-video, 10 seconds

A highland meadow where a small herd of cumulus clouds has descended to graze, drifting a meter above the grass, trailing thin wisps as they crop the turf bald in slow patches. Static telephoto wildlife shot, heat-haze shimmer, absolutely matter-of-fact as if this is normal behavior. Soft overcast light, muted greens. The only sound is wind, distant sheep bells, and a low woolly rumble whenever a cloud tears up grass. No music.

Battle robot, Image-to-video, 20 seconds

Create an ultra-photorealistic, ultra-cinematic sci-fi war sequence with major blockbuster quality. Use @Image1 ima as the visual inspiration for the main robot scale, mood and composition. The world must feel like real live-action footage. Show a lone human warrior standing in front of a colossal battle robot. He raises his arm and waves or signals toward the machine. In that instant, the giant robot powers on and begins to move with terrifying weight and realism. Its head, armor plates, pistons, cables and mechanical joints shift naturally. Every step makes the ground shake violently, sending dust, debris and loose metal trembling across the battlefield. Start with an intense front-facing shot, then pull the camera farther back to reveal the true scale of the war zone. As the camera widens, show that the lone warrior is only the front figure of a massive human army: rows of soldiers, armored fighters, vehicles and battle-ready units stretching across the terrain. Behind and around the lead machine, reveal an army of huge robots in different shapes and silhouettes, including towering walkers, heavy mech units, giant tank-like war machines and industrial combat machines preparing for battle. Add more movement as engines ignite, lights glow, steam vents burst and dirt scatters under their weight. Above them, huge dark sci-fi warplanes and threatening aerial craft roar into the sky, casting moving shadows over the battlefield. Add dramatic wind, smoke, dust clouds, distant explosions, sparks, atmospheric haze and powerful cinematic sound-driven energy. The mood is tense, epic and frightening, like the final buildup before an enormous war. Emphasize scale, weight, realism, camera shake, military formation, mechanical detail and high-end visual effects.

Soldier standing before towering white mech robots in a desolate, cloudy wasteland

Painting, image-to-video, 20 seconds

Create an ultra-photorealistic Hollywood-level museum transformation sequence. Use @Image1 im as the portrait reference. The painting hangs inside a grand classical museum in an ornate antique frame. Preserve the woman from @Image1: same face, age, body proportions, long wavy blonde hair, white historical dress and every clothing detail. Nothing about her identity or costume may change. Begin with a slow cinematic push toward the framed painting. At first it is still. Then tiny cracks of warm light glow beneath the painted surface. The oil-paint texture subtly ripples like canvas under pressure, never warping her face. Her eyes move first, then she slowly turns her head. One painted hand reaches forward and physically breaks through the flat surface into the museum. Fine pigment dust, golden particles and faint luminous paint strands drift from the frame while interactive light falls naturally across her hand and wall. She pushes farther through: shoulder, head, hair and torso emerge with correct depth and anatomy. As each part crosses the frame boundary, painted texture transforms seamlessly into real human skin, individual hair strands and real fabric, as though the painting is becoming physical. The transition remains perfectly aligned with @Image1. The canvas flexes in subtle waves around her body, then settles. Add restrained volumetric light, floating museum dust, delicate paint particles, soft lens bloom and slight atmospheric distortion around the frame, creating premium feature-film VFX.

What does Flux 3 cost on Aristotto?

Flux 3 is available on Aristotto with the 46% off on all paid tiers through December 20, 2026. Credit cost varies by clip length, and extension costs more per second than fresh generation, worth knowing before you plan a workflow built on extending short clips rather than generating longer ones from the start.

Common questions

What is Flux 3?

Flux 3 is Black Forest Labs' multimodal video generation model, trained on video, images, and audio in one unified architecture. It generates clips from 5 to 20 seconds with native synchronized audio.

Is Flux 3 available on Aristotto?

Yes, with the 46% per-generation credit reduction on all paid tiers through December 20, 2026.

How long can a Flux 3 clip be?

5 to 20 seconds in a single generation, any whole number of seconds. Clips can be extended using the video continuation workflow.

How is Flux 3 different from Seedance or Wan?

The core difference is architectural. Seedance and Wan generate video and add audio in separate steps. Flux 3 learns both together, so sound events are natively synchronized to visual events.

What is Flux 3 video-to-video?

FLUX 3 video-to-video (V2V is an AI generation technology by Black Forest Labs that transforms an existing video clip into a completely new visual style, animation, or environment based on a text prompt. Unlike standard video generators, it uses the structural composition and motion data of your original footage to ensure flawless continuity and realistic movement in the final output.

What are Flux 3 keyframes?

A workflow where you provide up to 10 ordered still images with specific frame positions, and the model generates continuous motion between them. Selecting a frame position also determines how many seconds each beat gets.

What is the Flux 3 HARD CUT syntax?

Write HARD CUT between timestamped ranges in a prompt to tell the model where an angle or scene change should happen within a single generation.

Does Flux 3 generate audio automatically?

Yes, in the same pass as the video, at no extra cost. Physical events in the scene produce the strongest audio results.

Can Flux 3 do different visual styles?

Yes, and this is one of its real strengths. It handles claymation, VHS found footage, comic-book halftone, documentary, and fully stylized work with genuine commitment, not just a filter over the same default look.

Does Flux 3 support 4K?

No, FLUX 3 generates at 720p and 1080p.

Discover more

View all