Aristotto
Back
Guides

How to Prompt GPT Image 2

Aristottoby Aristotto6 min
A boxer slumped in the corner between rounds while his trainer holds a towel and a cutman treats a cut above his eye.

GPT Image 2 launched April 21, 2026, and it behaves differently from the image models people are used to. It's built on a reasoning step, it plans the composition, works out object counts, lighting, and spatial relationships before generating a single pixel. That one architectural fact changes how you should write a prompt for it.

Structure over keyword stacking

Old habit: "cinematic lighting, 8K, masterpiece, trending on artstation." That does nothing here. GPT Image 2 reads descriptive natural language, not a pile of quality tags.

What actually works is a consistent structure: background and scene first, then subject, then key details, then constraints. Write it the way you'd brief a photographer or an art director, not the way you'd tag a stock photo.

Getting text right

This is the model's real strength. Put the exact wording you want in quotation marks, specify the font style, weight, color, and placement, and add "verbatim, no extra characters" when accuracy actually matters. Spell out tricky brand names letter by letter if the model keeps getting them wrong. Skip any of that and it will sometimes invent extra characters or repeat a headline that should only appear once.

Editing without losing what you already have

For edits, state what changes and what must stay untouched, explicitly, every time. A clean pattern that holds up:

Change: [exactly what should change]. Preserve: [face, identity, pose, lighting, framing, background, text, layout]. Constraints: [no extra objects, no logo drift, no watermark].

Models drift silently when you stop repeating the rules. If a previous edit kept the lighting warm and you don't restate that on the next pass, don't be surprised if it shifts.

For multi-image work, combining a person from one photo with an outfit from another, label each input by its role: "Image 1: base scene to preserve. Image 2: jacket reference." The model can take up to 16 reference images on edits, but it needs to know what each one is for.

Resolution and quality settings

Native output goes up to 2K (2560x1440), with an experimental 4K option that OpenAI itself flags as more variable. Default to 1K or the "low" quality setting while you're iterating, and only bump up to 2K or "high" for a final pass, higher settings cost more and take noticeably longer, sometimes three or more minutes at full 4K.

What to know before you rely on it

No transparency support, no RGBA output. If you need a transparent background, you'll need a different tool or a background-removal pass afterward. Brand logo reproduction is hit or miss on fine detail, treat generated logo concepts as a starting point, not a final asset. And the model isn't deterministic, the same prompt produces a different result each run, so budget for two or three regeneration passes rather than expecting to one-shot it.

Four prompts that work

Product photography

Background: a matte concrete surface with soft natural light from the upper left. Subject: a pair of leather boots, angled three-quarter view, laces slightly undone. Details: visible leather grain, subtle scuff marks for authenticity, brushed metal eyelets. Constraints: no text, no logos, photorealistic, no extra objects in frame.

White over-ear headphones with 'AURA' embossed on each cup, laid flat on a purple surface.

Infographic

Background: clean white background with a subtle grid. Subject: a step-by-step infographic titled "How Coffee Is Roasted" with four numbered stages: green beans, first crack, development, cooling. Details: simple line-art icons for each stage, connected by a horizontal arrow, muted brown and cream color palette. Constraints: EXACT TEXT for the title and each stage label, verbatim, no extra characters, clean sans-serif typography throughout.

Photorealistic documentary style

Background: a small family-run bakery early in the morning, flour dust visible in a shaft of window light. Subject: an older baker shaping dough by hand, focused expression, flour-dusted apron. Details: visible flour on the hands and forearms, natural skin texture, worn wooden work surface. Constraints: photorealistic, documentary photography style, no text, no staged expression, natural unposed moment.

A woman in a trench coat reaching for a subway door on a crowded platform at 42nd Street station.

UI mockup

Background: a clean smartphone app interface, light mode, minimal white background. Subject: a fitness tracking app home screen showing a circular progress ring, a step count, and three summary cards below. Details: "8,432 STEPS" as the primary large text, "Today's Goal: 10,000" as a smaller subtext, soft blue and white color scheme. Constraints: EXACT TEXT as specified, verbatim, clean modern UI style, no extra elements, realistic phone screen proportions.

Where creators are actually using this

Two things stand out about how people use GPT Image 2 in practice, and one of them pairs directly with Seedance.

Storyboarding before animating. A workflow shows up again and again across independent creator tutorials: build a storyboard in GPT Image 2 first, annotate camera moves and timing across the frames, then feed the strongest panels into Seedance to animate. One documented case involved a complex market scene with a moving crowd, going straight to video first failed five times in a row, chaotic movement, camera transitions that didn't track. Storyboarding it first, with every shot's timing and camera intent annotated, fixed it. The reason this pairing works: GPT Image 2 reasons about narrative and composes a scene's meaning across panels rather than just rendering a literal shot list, which is exactly what a storyboard needs.

Realism, with a caveat worth knowing. GPT Image 2 is widely regarded as a genuine step up on realism. One reviewer comparing it against Nano Banana 2 put it plainly: "Nano Banana 2's image looks rather cartoon-ish" by comparison. That said, a careful side-by-side test found it actually leans toward a "cleaned up," polished look rather than gritty documentary naturalism, faces read as tidy, environments as neat. If you want true grit rather than a polished default, push for it explicitly in the prompt.

Common questions

Why does my prompt keep producing a second, unwanted subject?

The model fills gaps you didn't specify. State explicitly what shouldn't appear, not just what should.

Does GPT Image 2 support transparent backgrounds?

No. Use a separate background-removal step if you need RGBA output.

Is 4K worth using for production work?

OpenAI flags anything above 2K as experimental. For reliable results, generate at 2K and upscale separately if you need a larger final size.

Discover more

View all