Aristotto
Back
Guides

Gemini Omni 1.1 Flash: Review and Prompting Guide

Aristottoby Aristotto12 min

Gemini Omni Flash 1.1 is Google DeepMind's multimodal video model, updated August 27, 2026. It generates 3 to 10 second clips with native synchronized audio, then keeps the take live in the conversation: a follow-up prompt edits that specific clip, and another extends it, up to 40 seconds total.

That conversational layer is what makes this model genuinely different from everything else on Aristotto. Every other video model hands you a clip and steps back. Gemini Omni Flash 1.1 stays in the room. Available now on Aristotto.

What is Gemini Omni Flash 1.1?

Gemini Omni Flash was announced at Google I/O on May 19, 2026, with the API opening June 30. The 1.1 update shipped August 27, 2026. Google frames it explicitly as "a new suite of creative controls and generative video capabilities," not a new base model. The underlying model card covers both versions together.

It generates clips from 3 to 10 seconds at 24 FPS, from 360p up to 4K, in 16:9 or 9:16. Every clip includes native audio produced in the same pass. Inputs are text, images, and video. Audio reference upload is architecturally supported but not yet live in the API.

Five task modes cover everything: text-to-video, image-to-video, reference-to-video, edit, and extend. The old 1.0 model ID (gemini-omni-flash-preview) deprecates September 30, 2026.

One thing worth knowing before you reach for 4K: independent testing confirmed that 1080p and 4K are upscales of the 720p render, not independent higher-quality generations. The pixel correlation between 720p, 1080p, and 4K frames from the same prompt sits above 0.998. Draft and approve at 720p, then output at 1080p or 4K at the delivery stage. The 360p draft mode, on the other hand, is a genuinely different render, useful for quick directional tests before committing to a full generation.

What actually changed in 1.1

Five things are new. Everything else stayed the same.

Scene extension to 40 seconds. The original Omni Flash capped hard at 10 seconds with no extension. 1.1 adds continuations of 3 to 10 seconds each, using up to 10 seconds of prior footage as context, up to a 40-second cumulative maximum.

First-and-last-frame control. You now specify exact opening and closing stills, and the model fills the motion between them. This is also how camera control works on this model: instead of a camera parameter, you pin where the shot starts and ends.

360p draft mode. Faster and cheaper than standard generation. Use it for directional tests, not for final previews, since independent testing shows 360p diverges meaningfully from what the 720p render will actually look like.

Free-form durations. 1.0 only accepted 3, 5, or 10 seconds exactly. 1.1 accepts any integer from 3 to 10, eliminating the waste of padding a 6-second shot to 10.

Working video references. Video reference input existed in 1.0 but didn't process correctly at launch. It works in 1.1: up to 3 clips of up to 3 seconds each, used for character or motion references. Audio inside video references is ignored.

What Google's own updated model card says did not improve: "maintaining complete consistency throughout edits, generating scenes with complex motion, or rendering perfectly accurate text remains a challenge." Those words are verbatim from the August 27, 2026 model card. Any article claiming 1.1 fixed consistency or text accuracy is not citing a Google source.

Where Gemini Omni Flash 1.1 is strongest

Conversational editing. Make one focused change per prompt, name what stays the same, and the model holds the rest of the shot intact. Face, wardrobe, camera angle, all persist. The edit lands on the thing you described and leaves everything else where it was. About four chained edits is the reliable window before drift starts appearing.

Building to 40 seconds without stitching. Each extension appends 3 to 10 seconds, reading the final moments as context. A 10-second take becomes 20, then 30, then 40, all from the same original take with no visible join. The extension timecode resets per segment, so "after 2s" in an extension means 2 seconds into the new footage, not into the original clip.

First-and-last-frame interpolation. For product work especially, this moves the iteration into an image model where it's cheaper and faster, then only pays for video once both endpoints are approved. Pass the same image as both first and last frame and the clip loops.

On-screen text in English. Signs, titles, lower thirds, and word-by-word kinetic text all render legibly when you specify exactly what each surface says. Define even background text, or the model invents its own. Non-Latin scripts are the documented weak spot, specifically Japanese and Chinese, where independent testing found the majority of characters rendering incorrectly.

Audio on demand. Dialogue comes back lip-synced in a voice that holds across edits and extensions. The audio track you describe in the prompt is the track you get. Silence has to be explicitly asked for.

The rule most guides get backwards

This is the single most misunderstood thing about Gemini Omni Flash 1.1. Google's own documentation states it directly:

A colon after a speaker's action means synthesized speech, rendered as audio. Quotation marks mean the text renders visibly on screen as subtitles burned into the frame.

A woman says: My name is Clara. → she speaks the line aloud, lip-synced.

A woman says: "My name is Clara." → the words appear as text in the frame, not spoken.

Get this backwards and your dialogue becomes subtitles. Every guide that doesn't mention this rule is missing the most practically important thing about prompting dialogue on this model.

How to write a prompt for Gemini Omni Flash 1.1

The model defaults to building a multi-shot narrative. If you want one continuous take, you must say so explicitly. Three phrases suppress the multi-shot behavior: "single continuous shot," "in a single unbroken scene," or "no scene cuts." Leave it unspecified and you'll usually get several cuts.

A prompt is a short shot brief in this order: subject and action, scene, camera and light, audio, shot structure. Aspect ratio, resolution, and duration are settings, not prompt text.

Describe audio every time. An unspecified audio prompt gets a model-invented soundtrack, usually generic music. "No music, just room tone" is the minimum. "A jazz piano two blocks away, rain, her footsteps on wet stone" is the specific version.

Event-based timing beats abstract timing. "When she touches the mirror, it ripples like water" lands more reliably than "at 5 seconds, the mirror ripples." Event-conditional instructions give the model something physical to wait for.

Use timecoded blocks for sequenced action. The model follows [0-3s], [3-7s], [7-10s] blocks inside a prompt. One beat per block, not multiple actions stacked into the same range.

For edits: one change at a time. State the change, then add "Keep everything else the same." That phrase is the only documented preserve instruction in Google's own materials, and it does real disambiguation work. Everything you describe in an edit prompt is something the model may re-render, so describe only the change.

Negatives go in the prompt body. There is no negative prompt field. Write "No dialogue," "No text overlay on screen," "No scene cuts" directly inside the main prompt.

Reference tags

Tags assign each uploaded asset a specific role:

<FIRST_FRAME> sets the exact opening frame. <LAST_FRAME> sets the closing frame, always paired with a first frame. Pass the same image as both and the clip loops. <IMAGE_REF_0> through <IMAGE_REF_9> mark reference images by upload order, each needing an explicit job stated in the prompt: "the person from <IMAGE_REF_0> walks through the market." <VIDEO_REF_0> marks a reference clip for character or motion identity, kept under 3 seconds.

Attach all references before the first generation. Adding a reference mid-conversation destabilizes a shot that was holding. Reference images and first/last frames cannot be combined in the same call.

Five prompts across different workflows

The location reveal, Text-to-video, 10 seconds

[0-3s] A slow push-in on a closed wooden door at the end of a narrow hallway, warm light leaking under the door, no sound yet. [3-7s] The door swings open slowly inward, revealing a vast library with cathedral ceilings, thousands of books, afternoon sunlight from a high window. [7-10s] The camera continues through the doorway and floats slowly toward the centre of the room. Single continuous shot, no scene cuts. Ambient sound: the creak of the door hinge, the soft breath of air as it opens, distant birdsong from the window, no music.

The relighting edit chain, Edit

Generate the first clip as text-to-video, then use these follow-ups in sequence:

Generate: A woman sits reading at a small wooden desk by a window, morning light soft and even through the glass, a cup of tea steaming beside her. Static medium shot, no movement. Single continuous shot. Room tone only, an occasional page turn.

Edit 1: Change the lighting to late evening. The window shifts from bright morning to deep blue dusk, a warm lamp on the desk fills the room with a softer glow, shadows deepen on the left wall. Keep everything else the same.

Edit 2: Add a cat curled on the corner of the desk, sleeping. Keep the woman, the lighting from the last edit, the desk, and the camera angle exactly as they are.

The product reveal loop, First-and-last-frame, 8 seconds

Start frame: a glass perfume bottle centered on dark marble, studio light from the left, lid closed.
End frame: identical composition, lid open, a faint mist of fragrance visible above the opening.

The bottle stays perfectly still as the studio lights shift: overhead softbox dims, a warm rim light rises from behind the glass, the amber liquid glows from within. The lid lifts slowly in the final two seconds. One continuous take, no cuts, no camera move. A faint mechanical click as the lid releases, the soft hiss of fragrance dispersing, quiet room tone, no music.

The singing character, Reference-to-video, 9 seconds

<IMAGE_REF_0> performs on a small stage under a single warm spotlight, seated on a stool, looking slightly down at the strings of an acoustic guitar for the first four seconds, then lifting her gaze toward the audience for the final five. Preserve her hair, jacket, and facial features from the reference exactly. Camera holds on a medium shot, slow push-in beginning at 5 seconds. She says: My whole life I've been waiting for this moment. No other music behind the line. Keep everything else the same.

The scene extension, Extend

Extend this video. The scene continues: she closes the book, sets it down on the desk, and stands. She crosses to the window and looks out at the street below as the lamp casts her silhouette against the glass. The camera holds its position in the room and does not follow. Street noise muffled by the glass, her quiet exhale, the lamp hum underneath.

What it costs on Aristotto

Gemini Omni Flash 1.1 is available on Aristotto. Cost varies by clip length and resolution. Draft at 720p. That is the resolution the model actually renders at, and independent testing confirmed that 1080p and 4K deliver the same underlying render at larger sizes. Upscale at the delivery stage if the job requires it.

Common questions

What is Gemini Omni Flash 1.1?

Google DeepMind's multimodal video model, updated August 27, 2026. It generates 3 to 10 second clips with native audio, and the take stays live in the conversation for editing and extension up to 40 seconds total.

Is Gemini Omni Flash 1.1 available on Aristotto?

Yes, Gemini Omni Flash 1.1 is available on Aristotto.

How long can a Gemini Omni Flash 1.1 clip be?

Each generation produces 3 to 10 seconds. Extend up to a 40-second cumulative maximum through follow-up prompts, each adding 3 to 10 seconds.

How does Gemini Omni Flash 1.1 conversational editing work?

After generating a take, describe one specific change in a follow-up prompt and add "Keep everything else the same." The model applies the change while preserving everything you didn't mention.

How is Gemini Omni Flash 1.1 different from 1.0?

1.1 adds scene extension to 40 seconds, first-and-last-frame control, 360p draft mode, free-form durations from 3 to 10 seconds, and working video references. The original model ID deprecates September 30, 2026.

Should dialogue go in quotation marks?

No. Use a colon after the speaker's action for spoken dialogue. Quotation marks tell the model to render the text visibly on screen instead of speaking it.

What is Gemini Omni Flash 1.1 first-and-last-frame?

A workflow where you pin exact opening and closing stills, and the model generates the motion between them. Pass the same image as both and the clip loops.

Is the Gemini Omni Flash 1.1 4K output actually 4K?

It's a 4K file, but independent testing confirmed it's an upscale of the 720p render. The pixel correlation between 720p and 4K frames from the same prompt is above 0.998.

How many edits can I chain on one take?

About four before drift appears. Past that, re-anchor with "Keep all character details exactly as they are, only change X."

What is Gemini Omni Flash 1.1's biggest weakness?

Google's own model card lists three: consistency through edits, complex motion sequences, and accurate text rendering. These are unchanged from 1.0.

Does Gemini Omni Flash 1.1 multi-shot by default?

Yes. It builds a narrative across several shots unless you add "single continuous shot, no scene cuts" to the prompt.

Discover more

View all