ManyMotions
Media ForgeGenerate images & video, fastMedia StudioNewBuild images & video by chatting with AIWorkflowsAutomate pipelines at scaleClipperNewBreak a video down & cut the highlightsMarketing AgentDone-for-you UGC video ads
MCPPricingFAQ
Log inStart for $1
AI Models/GPT Image/Prompting

How to prompt GPT Image

GPT Image reads a prompt like a creative brief, not a keyword list. Write it in a fixed order — scene, subject, key details, then constraints — and tell it what the image is FOR ("a product ad", "a UI mockup", "an infographic"): the stated use case sets the polish and layout mode. Exact text goes in quotes and renders verbatim, plain-language exclusions ("no extra elements") are respected, and editing works as a conversation — state the one change, then the list of everything that must not move.

Anatomy of a GPT Image prompt

  1. 1

    Scene & background

    Open with where the image lives: "a bright kitchen counter at morning", "a clean studio sweep in pale gray". The model builds the stage first, then places everything on it.

  2. 2

    Subject

    The main subject with concrete attributes: materials, colors, condition — "a glass bottle of cold-brew with a kraft-paper label". Specifics here are what the model preserves through iterations.

  3. 3

    Key details

    Framing, perspective and lighting in plain terms: "close-up, eye-level, soft diffuse light, logo top-right, negative space on the left". Explicit placement instructions are followed literally.

  4. 4

    Text to render

    The exact copy in quotes, plus typography: font style, weight, color, placement. For tricky brand names, spell them letter by letter. Ask for it "exact, no extra characters".

  5. 5

    Intended use

    One phrase naming the deliverable — "this is a social ad", "an onboarding screen mockup", "a printed infographic". It switches the model into the right layout and finish conventions.

  6. 6

    Constraints

    What must appear, what must not, and — in edits — the full preserve list: "change only the background; keep the face, pose, lighting and label exactly the same". Anything not locked is fair game.

Template

[scene/background], [subject with materials and colors], [framing, perspective, lighting, placement], text: "EXACT COPY" in [font style, weight, color, placement], intended use: [ad / UI mockup / infographic / packshot], constraints: [what must stay exact, what must not appear]

10 example prompts that work

Social ad with a verbatim headline

GPT Image 2 · 2K · 9:16 · quality high
A vertical social ad. Scene: deep navy studio background with a soft radial glow. Subject: a matte-white wireless earbuds case, open, floating at a slight angle, crisp reflection below. Headline text: "HEAR EVERYTHING." in a bold geometric sans-serif, white, centered in the top third. Smaller text below the product: "Wireless. 40-hour battery." in light gray. Intended use: paid social ad. Constraints: exact text as written, no extra characters, no other text anywhere.

Quoted copy plus a font description plus "no extra characters" is the verbatim-text recipe. Labeled segments (Scene / Subject / Text / Constraints) keep a dense brief unambiguous — the model follows the structure, not just the words.

Infographic with labeled sections

GPT Image 2 · 2K · 3:4 · quality high
A clean vertical infographic titled "HOW COLD BREW IS MADE" in bold dark-blue letters at the top. Four numbered sections stacked below, each with a simple flat icon and one short label: 1. "Coarse grind", 2. "Steep 18 hours", 3. "Filter twice", 4. "Serve over ice". White background, generous white space, consistent 2-color palette of dark blue and warm orange, flat vector style. All text exactly as written, readable and horizontal.

GPT Image handles structured layouts — numbered sections, icon plus label pairs, consistent spacing — that keyword-style models scramble. Dense text layouts need quality high; every label is quoted so nothing gets paraphrased.

Photoreal product shot

GPT Image 2 · 2K · 1:1 · quality high
A real photograph, morning window light. Scene: a pale oak café table near a window. Subject: an iced matcha latte in a tall ribbed glass, distinct green-to-white gradient, condensation beading on the glass, a metal straw. Key details: close-up at eye level, shallow depth of field, blurred café interior behind, natural soft shadows. Intended use: menu photography. Constraints: no people, no text, no logos.

"A real photograph" plus camera-adjacent language (eye level, shallow depth of field) engages the realism mode more reliably than stacking "photorealistic 8K". The constraint line cleanly excludes the usual drift: stray text and invented branding.

UI mockup that looks shipped

GPT Image 2 · 2K · 16:9 · quality high
A desktop dashboard UI for a habit-tracking app, presented as a straight-on screenshot. Layout: left sidebar with the app name "Streaks" and five nav items ("Today", "Habits", "Stats", "Friends", "Settings"), main area with a weekly grid of colored completion dots and a large card reading "12-day streak". Light theme, soft gray background, one accent color: teal. Realistic interface spacing and hierarchy, real UI text only, no lorem ipsum, no concept-art flourishes.

Describe the product as if it already shipped — real nav labels, real hierarchy — and ban concept-art language explicitly. Quoting every UI string keeps labels crisp and prevents invented menu items.

Packaging with dense label copy

GPT Image 2 · 2K · 3:4 · quality high
A photoreal packshot of a cylindrical tea tin on a white sweep. Front label layout, top to bottom: brand name "KOYAMA" (spell it exactly: K-O-Y-A-M-A) in a serif, dark green; below it "Roasted Hojicha" in smaller caps; a thin gold rule; then "Loose leaf - 80g" at the base. Soft even studio lighting, slight top-down angle. Constraints: all label text exact and correctly spelled, no additional words on the label.

Letter-by-letter spelling is the documented fix for brand names the model wants to "correct". Describing the label as an ordered layout (top to bottom) gets typography hierarchy right on the first pass.

Sticker asset on a transparent background

GPT Image 1.5 · 1K · 1:1 · background transparent · PNG
A die-cut sticker illustration of a smiling cartoon avocado wearing sunglasses, thick white sticker border, bold flat colors with simple cel shading, slight drop shadow inside the sticker edge. Transparent background, nothing outside the sticker outline.

GPT Image 1.5 exposes a background control — set it to transparent and export PNG for a drop-in asset. The "nothing outside the outline" constraint keeps the alpha clean of stray marks.

Edit: change one thing, preserve the rest

GPT Image 2 Edit · 2K · match source ratio · quality high
Change only the wall color behind the sofa to sage green. Keep everything else exactly the same: the sofa, cushions, rug, plant, window light, shadows, and framing must be identical to the original photo.

The change-plus-preserve contract is the core editing pattern: one sentence for the change, one for the lock list. Anything you leave unlisted the model may reinterpret — the preserve list is not politeness, it is the spec.

Edit: virtual try-on with identity lock

GPT Image 2 Edit · 2K · match source ratio · quality high
Replace the man's gray t-shirt with a navy chunky-knit crewneck sweater. Do not change his face, facial features, skin tone, hair, body shape, pose, or identity in any way. Keep the background, camera angle and lighting identical, and make the sweater fabric fold and drape realistically with the existing light.

Identity-sensitive edits need the explicit lock phrasing — face, features, skin tone, pose, identity — repeated on every iteration to prevent drift. Asking for realistic fabric behavior under the existing light is what stops the garment looking pasted on.

Edit: composite two input images

GPT Image 2 Edit · 2K · quality high · 2 input images
Image 1: a product photo of a green glass skincare bottle. Image 2: a bathroom shelf scene with morning light. Place the bottle from Image 1 on the shelf in Image 2, centered on the middle shelf, matching the scene's lighting, adding a soft natural shadow under the bottle. Keep the bottle's label, proportions and color exactly as in Image 1; keep the scene otherwise unchanged.

With multiple input images, refer to each by index and state the interaction explicitly — which element travels, where it lands, whose lighting wins. Locking the label and proportions keeps the product legally accurate.

Edit: iterative micro-refinement

GPT Image 2 Edit · 2K · match source ratio
Make the lighting slightly warmer, like late golden hour. Change nothing else — composition, subject, colors of the objects, and all text stay exactly as they are.

The documented refinement loop: get a clean base, then one small change per pass instead of piling adjustments into a single prompt. Each micro-edit re-states the freeze so successive passes don't accumulate drift.

Settings that matter

  • Quality

    low / medium / high, default high. Use low to explore compositions fast and in volume; high is non-negotiable for dense text, infographics, close-up faces and anything a client will zoom into.

  • Output format

    PNG (default) for graphics, UI and anything with text or transparency; JPEG for photographic outputs where file size matters; WebP for direct web use.

  • Background

    GPT Image 1.5 and 1 only: auto / transparent / opaque. Set transparent (and keep PNG) for stickers, logos and cut-out product assets.

  • Resolution

    GPT Image 2 runs 1K, 2K and 4K — draft at 1K, ship at 2K, reserve 4K for print and dense layouts. GPT Image 1.5 and 1 output at 1K.

  • Aspect ratio

    GPT Image 2 adds 16:9 and 9:16 to the square and 4:3/3:4 set — the pick for widescreen mockups and vertical ads. The 1.5 and 1 generations cover 1:1, 3:4 and 4:3.

  • Images per run

    Generation returns up to 4 images per run — use the full batch at quality low to shortlist a composition, then rerun the winner at high. Edits return a single result.

Do and don't

Do

  • Keep one section order in every prompt — scene, subject, key details, constraints — and reuse it; consistency compounds.
  • Break complex briefs into short labeled segments or line breaks instead of one long paragraph.
  • Name the intended use ("paid social ad", "UI mockup", "infographic") — it sets layout and finish conventions.
  • Put every piece of literal text in quotes, specify font style, weight, color and placement, and demand "exact, no extra characters".
  • Split every edit into change + preserve, and repeat the preserve list on each iteration.
  • Refine iteratively: one small change per pass ("make the lighting warmer") beats one overloaded prompt.
  • With multiple input images, reference each by index ("Image 1", "Image 2") and describe how they interact.

Don't

  • Don't spam quality keywords — "8K, masterpiece, trending" does nothing here; the model reads descriptive language, not tags.
  • Don't leave the preserve side implicit in an edit — anything unlisted is treated as editable.
  • Don't stack five changes into one edit prompt; drift multiplies with each simultaneous change.
  • Don't request small or dense text at quality low — it renders soft and approximate.
  • Don't describe a UI mockup in concept-art language ("futuristic sleek design") — describe the shipped product with real labels.
  • Don't paraphrase your own headline between iterations; repeat the quoted string exactly or the model will "improve" it.

Advanced techniques

The text engine: verbatim copy in images

Text rendering is this family's signature capability, and it is prompt-controllable to a degree unusual for image models. The working recipe: quote the exact string, give it a typographic spec (font style, weight, color), pin its placement relative to other elements, and close with "exact, no extra characters". For words the model wants to autocorrect — invented brand names, foreign spellings — spell them out letter by letter. Long copy works too: multi-line labels, numbered lists and paragraph blocks hold up at quality high, which is where dense text belongs.

Editing as a conversation

GPT Image editing is instruction-driven: you describe the delta in plain language and contract the rest with a preserve list. The pattern scales from one-word recolors to full try-on and compositing work, and it iterates — each pass takes the previous output plus a new single change. Two rules keep long chains stable: repeat the preservation constraints every time (they are per-request, not remembered), and keep each step small. Identity is the most drift-prone attribute, so face, features, pose and skin tone get named explicitly whenever a person is in frame.

  • Recolor / swap: "Change only X to Y; keep everything else the same"
  • Identity lock: "Do not change her face, facial features, skin tone, body shape, pose, or identity in any way"
  • Relight: name the target light and freeze the rest — "warmer, like late golden hour; change nothing else"
  • Style transfer: state what changes (medium, palette) and what holds (composition, subject), plus "no extra elements"

Multi-image compositing by index

The edit models accept several input images at once, and the model resolves them best when your prompt treats them like named layers: "Image 1: the product. Image 2: the scene. Put the product from Image 1 on the counter in Image 2." Say which image wins on style, whose lighting applies, and what must survive untouched — typically the product's label, colors and proportions. This covers style-reference workflows too: "apply Image 2's illustration style to the subject of Image 1" is understood as an instruction, not a collage.

World knowledge and the use-case switch

Because the family reasons over the request before rendering, it applies real-world knowledge to layout problems: an infographic gets readable hierarchy, a UI mockup gets plausible spacing and controls, a diagram gets sensible flow — from a description of the content alone. Steer this by declaring the deliverable ("a printed infographic explaining X", "an onboarding screen for a banking app") rather than describing pixels. If a structured brief suits you better, labeled segments or a JSON-like block of fields work as well as prose; the model cares that intent and constraints are unambiguous, not which format carries them.

Which variant to use

  • GPT Image 2

    Default for everything: highest fidelity, best text rendering, 1K-4K output, widescreen 16:9 and vertical 9:16, up to 4 images per run.

  • GPT Image 2 Edit

    The editing counterpart — change-plus-preserve edits and multi-image compositing at up to 4K.

  • GPT Image 1.5

    When you need the background control: transparent-background assets (stickers, logos, cut-outs) at 1K in square or 3:4/4:3.

  • GPT Image 1.5 Edit

    Composition-preserving edits at 1K — fine for social-size iterations that don't need 2K/4K output.

  • GPT Image 1

    The original generation with solid text rendering and an auto quality mode; keep for continuity with existing prompt sets.

Frequently asked

How do I get text rendered exactly right?

Quote the exact string, describe font style, weight, color and placement, add "exact, no extra characters", and generate at quality high. Spell fragile words letter by letter. GPT Image 2 sustains this even on dense, multi-line layouts.

Is there a negative prompt?

No parameter — and none needed. Plain-language exclusions in the constraints line are followed: "no people, no text, no logos, no extra elements". State them as part of the brief rather than as a separate syntax.

How do I edit a photo without the rest changing?

Use the change-plus-preserve pattern: "Change only X. Keep everything else exactly the same: [list]." The list is load-bearing — unlisted elements are considered editable, and it must be repeated on every iteration.

How do I get a transparent background?

Generate on GPT Image 1.5 (or GPT Image 1) with background set to transparent and PNG as the output format, and add "nothing outside the subject outline" so the alpha stays clean.

Which quality setting should I use?

Low for exploring compositions in batches of 4, high for anything with text, faces or client eyes on it. The prompt transfers unchanged between quality levels, so drafting low and finishing high costs nothing in rework.

Can I combine my product photo with a separate scene?

Yes — GPT Image 2 Edit takes multiple input images. Reference them by index ("place the bottle from Image 1 into the scene in Image 2"), say whose lighting applies, and lock the product's label and proportions so it stays accurate.

Try GPT Image in Media Forge →About GPT ImageAll models
ManyMotions

Tell me what you need and grab a coffee. I'll have it ready before you finish the cup.

Product

Marketing AgentWorkflow StudioMedia ForgeTemplatesAI Models

Open in app

Workflow StudioMedia ForgeProjectsAvatars

Company

ContactPrivacy PolicyTerms of ServiceContent Moderation PolicyLegal NoticeCancel subscriptionWithdraw from contract

© 2026 ManyMotions

support@manymotions.com