Aristotto
Announcement

MiniMax H3 Is now on Aristotto

MiniMax H3 is live: 2K video with native stereo audio, sharp in-video text, and reference-driven editing from images, video, and audio.

MiniMax H3 doesn't just generate video, it understands text, images, video, and audio together as one context, so everything from the visuals to the dialogue, sound effects, and music comes out of a single generation instead of being stitched together afterward. It's especially strong with anything that needs to actually be readable in-frame: signage, subtitles, brand marks, UI mockups, game menus. It also holds up well in stylized and animated work, claymation, anime, and 3D character pieces, keeping a character's design consistent across shots when you give it a reference image to anchor to. Feed it up to 9 reference images, 3 video clips, and 3 audio clips at once (Omni Reference mode) to hold a character, style, or voice consistent across a whole sequence. And if one detail is off, you can send a localized edit instead of regenerating the whole clip. Prompt tip: name the camera move directly, like "slow orbit right" or "handheld tracking shot," rather than just describing the scene. H3 follows shot language closely, so naming the move gives you a more consistent result. Not sure what to test? "A claymation fox in a patchwork scarf sprints across a stop-motion autumn forest, leaves flying up behind each step. Handheld tracking shot at ground level, keeping pace with the fox's stride, slight camera bounce on impact. Crunchy leaf foley and a playful plucked-string score, camera whip-pans to follow as the fox darts behind a mossy log." "A sci-fi game HUD boots up over a pilot's cockpit view, radar ping and health bar sliding into frame with crisp, legible readouts. Slow push-in toward the main viewport as a targeting reticle locks onto an approaching ship, interface text sharpening as the camera settles. Low synth hum and a soft confirmation beep as the lock completes."