Aristotto
Back
Guides

How to Prompt Nano Banana 2

Aristottoby Aristotto5 min
Close-up portrait of a man laughing candidly outdoors, head tilted back with eyes crinkled.

Nano Banana 2 launched February 26, 2026, and it's not a pure diffusion model. It runs on Gemini's reasoning backbone, which means it plans the composition, resolves physics, and reasons about spatial relationships before it produces pixels. It also has something almost no other image model has: real-time web search grounding, it can pull actual facts and reference images from the web mid-generation.

The basic formula, then build up

Google's own starting point: "Create an image of [subject] [action] [scene]." Start there, then layer in specifics. The model rewards detail more than it rewards a clever structure, "a young woman in a red dress running through a park" beats "a woman in a red dress" every time.

Tagging references correctly

If you're using a reference image, tag it explicitly with @image1, @image2, and so on, and assign each one a name the model can follow through the scene. This isn't optional, skip the tag and the prompt frequently just doesn't work the way you expect. Nano Banana 2 accepts up to 14 reference images in a single prompt and can hold up to five consistent characters across a scene, but only if each one is clearly labeled.

What the web grounding actually changes

Because it can retrieve real information before generating, you can prompt it differently for anything fact-based: a real building's actual layout, a specific location's real appearance, current weather-style data for an infographic. The formula is: source or search request, then the analytical task, then how to visualize it. This is the thing that separates Nano Banana 2 from a model that's just guessing based on training data, ask it for the actual courtyard of a real building and it can go check rather than invent something generic.

Editing in plain language

Semantic editing means you describe the change, not the whole scene again: "change the sunny day to a rainy night" or "remove the person in the background." The model identifies what needs to shift and leaves the rest, lighting and reflections included, untouched.

Resolution, aspect ratio, and cost

Output ranges from a compact 512px option up through 1K, 2K, and 4K. It supports an unusually wide set of aspect ratios, from standard 16:9 and 9:16 through extreme formats like 1:8 and 8:1 for banners. If you're doing high volume iteration, the 512px setting keeps cost close to the original Nano Banana while you find the right composition, then upscale the final pick.

What to avoid

Same lesson as most current image models: keyword stacking like "4K, trending on artstation, masterpiece" adds nothing, the model understands natural language and ignores decorative spam. Write what you actually want in plain sentences instead.

Four prompts that work

Real-world grounded scene

Create an image of the actual courtyard of the Louvre Museum's glass pyramid at golden hour, accurate architectural proportions and glass panel layout, a small crowd of visitors in soft silhouette, warm evening light reflecting off the glass. Photorealistic, wide-angle perspective.

Character consistency across a scene

Using @image1 as the reference for the character's face and hairstyle, create an image of her sitting at an outdoor cafe table reading a book, soft afternoon light, a cup of coffee beside her. Maintain the exact facial features and hairstyle from @image1. Aspect ratio 4:5.

Semantic edit

Using the uploaded image, change the sunny afternoon lighting to a rainy evening, add wet reflections on the ground and light rain in the air, keep the subject's pose, outfit, and position in the frame completely unchanged.

Data-grounded infographic

Create a clean infographic showing average rainfall by month for Seattle, Washington, using accurate real-world data. Simple bar chart format, months labeled along the bottom, muted blue color palette, a small disclaimer at the bottom reading "Approximate historical averages."

Where creators are actually using this

The predecessor model's launch is one of the more well-documented viral moments in this space. Per Reuters, the original Nano Banana pulled 13 million first-time users into the Gemini app in just four days, and had generated more than 5 billion images by mid-October. Google leaned into the moment publicly, Sundar Pichai himself posted about it with a string of banana emoji.

Nano Banana 2 launched already embedded rather than standalone: it's the default across the Gemini app, Google Search's AI Mode and Lens in 141 countries, Google Ads, Flow, and the Gemini API. That distribution is worth knowing, if you're generating something meant to look native to a specific surface, a lot of your actual audience is going to encounter this model's output inside Search results and ads before they ever open a dedicated image tool.

Common questions

Why isn't my reference image being used correctly?

Almost always a missing @image tag. Tag every reference explicitly and name what it controls.

What's the difference between Nano Banana 2 and Nano Banana Pro?

Nano Banana 2 covers roughly 95 percent of Pro's capability at a fraction of the cost and speed. Reach for Pro only when 2 consistently fails a highly complex, multi-layered prompt.

Can it generate accurate real-world locations or data?

Yes, that's its main differentiator. It can pull real reference information from the web before generating, rather than guessing from training data alone.

Discover more

View all