Skip to main content
Framesail AI
All posts

GPT Image 2.5 vs GPT Image 2: 123 Prompts, Side by Side

GPT Image 2.5 Flare vs GPT Image 2 on 123 identical prompts and references: what got sharper, what the new model invents, and how to prompt it. With images.

By Jordan · Cofounder, Framesail

GPT Image 2.5 vs GPT Image 2 on the same anime cafe prompt and the same claymation river prompt, side by side

OpenAI released GPT Image 2.5 on September 8, 2026, as two API models: Flare, the fast one, and Sunburst, the precise-editing one. We took 123 images that GPT Image 2 had already rendered for us, replayed the exact same prompt, reference images, quality tier, and canvas size on GPT Image 2.5 Flare, and put every pair next to each other.

The short version: GPT Image 2.5 Flare renders in about 11 seconds where GPT Image 2 took about 18, texture and lighting are visibly richer, and it follows the written prompt more literally than the old model did. That last part cuts both ways. When the text and a reference image disagree, Flare sides with the text.

What OpenAI shipped

The ChatGPT Images 2.5 announcement describes sharper detail, more natural lighting, better subject preservation from reference photos, and up to 50% lower latency than Images 2.0. For the API there are two models. GPT Image 2.5 Flare is positioned as the default for most work. GPT Image 2.5 Sunburst is aimed at edit-heavy workflows where precision matters most.

Both support the generations and edits endpoints, and both accept the same quality settings as before plus two new tiers above high. On OpenAI's pricing page, the per-token rates for text input, image input, and image output are identical to GPT Image 2. Every pair below was rendered at the low quality tier at 2048 by 1152.

How we ran the comparison

The test is a replay, not a fresh prompt. Every image we generate is logged with the full payload sent to the model: the final prompt text, the reference images in order, the quality setting, and the output size. For each of 123 images we pulled that payload, changed only the model id from gpt-image-2 to gpt-image-2.5-flare, and rendered it again.

  • Five projects, five art styles: photoreal, claymation, anime, flat vector, and a stick-figure explainer. The photoreal set is a client project, so those pairs are described but not shown.
  • Three kinds of image: character and object references, four-view environment sheets, and shots from the first minute of each video, most of which carry two to four reference images.
  • One render per pair. Where a result looked like it could be noise, we re-ran that shot three times per model. Those cases are marked below.
  • Timing came from a separate run of six identical payloads on both models, back to back, so queue time is not in the number.

The left image in every pair is the original GPT Image 2 render. The right is Flare.

What got better

Texture and lighting. Clay shows thumbprints, seams, and tool marks. Paper shows grain. Metal picks up specular edges. Contrast and saturation sit a notch higher across all five styles. On claymation and photoreal this is the biggest single upgrade.

GPT Image 2.5 vs GPT Image 2 texture detail on a clay water drop reference and a shot of clay spheres

Speed. On six identical payloads run back to back, GPT Image 2 averaged about 18 seconds per image and Flare about 11. OpenAI's 50% claim holds up at the low quality tier.

Reference identity. Characters rendered from a reference image are the same person on both models. In one case Flare was better: a stick-figure shot named a specific character scratching his head, GPT Image 2 drew a generic figure, and Flare drew the right one. Where the prompt asked for tiny stick figures in sun hats, GPT Image 2 drew colored cartoon people and Flare drew stick figures.

GPT Image 2.5 drew the named character where GPT Image 2 substituted a generic stick figure

Multi-view environment sheets. We render environments as a four-view sheet: one place from four angles, which downstream shots use as a spatial anchor. Flare's sheets show one place. GPT Image 2's sheets drift between cells, sometimes into two different bridges or two different rivers.

Four-view environment reference sheets on GPT Image 2.5 vs GPT Image 2 for a suspension bridge and a river

Where GPT Image 2.5 reads the text harder than the image

This is the finding that changed how we prompt. GPT Image 2.5 weights every sentence it is given as an instruction, and when two inputs disagree, the text wins.

The style paragraph beats the shot. Each anime shot below was rendered with the same reference frames and the same shot description. The style paragraph, which sits above the shot text, asks for luminous skies, lens flares, and bloom. GPT Image 2 rendered the whole cafe sequence overcast and rainy. Flare rendered every shot sunny with a flare.

GPT Image 2.5 follows the style text over the reference frames, turning rainy anime shots sunny

The text beats the previous frame. The clearest case is a continuation shot: the model is handed the previous frame and told to keep its composition and change only what the new text describes. The frame had a mountain and a village in it. The text, which had gone stale after an edit, said the character was "on white." We ran it three times per model.

GPT Image 2 kept the mountain from the previous frame in all three trials; GPT Image 2.5 followed the text and dropped it

GPT Image 2 kept the mountain and the number scale in four out of four renders. Flare dropped the mountain in all three trials and kept the scale in one. Correcting the text to name what carries over fixed it in three out of three. Strengthening the instruction that the previous frame should win changed nothing. The lesson is not that Flare handles continuation badly. It is that Flare does what the sentence says, so the sentence has to be right.

It fills in what a description leaves open. An object reference prompt that said "no text, no additional props" produced a sealed envelope on GPT Image 2 and an envelope plus a letterhead page with readable text on Flare. A character described as a scientist in his 60s with grey hair and a beard came back with glasses and a tie. A damper cylinder was redrawn as a different mechanism. None of these are wrong on their own. They are decisions the old model did not make.

GPT Image 2.5 adds details a description leaves open: a letterhead, glasses and a tie, a redesigned damper

We tested the obvious fix in reverse: remove the "no text, no props" rule and let the model decide. It got worse. Without the rule, Flare pulled props from the style paragraph into the object shot and staged narrative context from the description. The rule stays. What matters more is that the description itself says only what the camera sees.

Framing: a label is not a framing

Eleven of the twelve close-up shots in our set opened with the word "Close-up." and nothing about how large the subject should be. GPT Image 2 filled the frame. Flare pulled back to a mid shot on every one of them. The single close-up that described its framing concretely, "fills the frame from the chest up," came back tight on both models.

We confirmed it with two claymation shots, two trials each. The label gave a mid shot four times out of four. Replacing it with "fills the frame from top edge to bottom edge" gave a close-up four times out of four.

GPT Image 2.5 renders a mid shot for the label Close-up and a true close-up when the prompt states the subject's size in frame

The same tendency shows up in diagrams. A flat, end-on schematic on a white background came back as a three-quarter perspective with piers. A vortex-shedding diagram gained an arrow-in stream and a pier and reads better for it. A shot of vibrating cables gained speed marks the flat style never asked for.

Flat vector diagrams on GPT Image 2.5 vs GPT Image 2: a schematic turned into a 3D view, a clearer vortex diagram, and added speed marks

How to prompt GPT Image 2.5

Everything above reduces to one rule. Say the visible thing, and only the visible thing.

  1. Describe the frame, not the label. "Her face fills the frame from chin to hairline" works. "Close-up" does not.
  2. Describe objects by what a camera sees. Purpose, symbolism, and story context ("being handed over," "hinting at a secret") get rendered as props and text.
  3. Never defer text. "Displaying some information" becomes a plausible fake name and a logo. Quote the exact words, or say nothing about text.
  4. For a shot that continues from a previous frame, name what carries over before naming what changes. If the text contradicts the frame, the text wins.
  5. Keep the style paragraph honest. If it lists a mood, a lighting effect, or a set of props, expect to see them in every shot.

Keep the negative constraints you already have. Removing them made things worse in our test, and they cost nothing.

Should you switch?

For stills that feed a video pipeline, yes, with the prompt discipline above. The speed gain is real, the texture gain is real, and reference identity holds as well as it did. The two things to check before switching a live workflow are close-up framing and any prompt that leans on a reference image to carry information the text does not state.

For text-heavy cards and literal diagrams, run a test first. Flare designs the card instead of transcribing it: a ten-item list became six with a rewritten heading, and titles doubled in size. That is a better designer and a less faithful copyist.

FAQ

What is the difference between GPT Image 2.5 Flare and Sunburst?

Flare is the fast model OpenAI positions as the default for most generation work. Sunburst is built for workflows where editing precision matters most, such as multi-turn edits on a photo. Both use the generations and edits endpoints and the same quality tiers. This comparison covers Flare only.

Is GPT Image 2.5 more expensive than GPT Image 2?

No. On OpenAI's pricing page the per-token rates for text input, image input, and image output are identical for gpt-image-2, gpt-image-2.5-flare, and gpt-image-2.5-sunburst. In our runs the token counts per image were also the same, so a render cost the same on both models.

How much faster is GPT Image 2.5 Flare?

On six identical payloads run back to back at the low quality tier, GPT Image 2 averaged about 18 seconds per image and Flare about 11. OpenAI claims up to 50% lower latency than Images 2.0, which matches what we measured.

Does GPT Image 2.5 keep characters consistent from a reference image?

Yes, as well as GPT Image 2 did and in some shots better. In our set, characters rendered from a reference image were recognizably the same person on both models across photoreal, anime, and stick-figure styles. Where Flare differs is on references generated from text alone, where it adds details the description leaves open.

Why does GPT Image 2.5 ignore my reference image?

It is not ignoring it. When the written prompt and the reference disagree, Flare follows the prompt. If a continuation shot says "on white" while the previous frame has a mountain, you get white. Fix the text so it names what should carry over, and the reference is honored again.

Why do my close-ups come out as mid shots on GPT Image 2.5?

Because the prompt says "close-up" and nothing else about size. Flare treats the label loosely. State how much of the frame the subject fills and what touches the edges, and the framing lands.

How we use it

framesail turns a script into a finished video: it writes the storyboard, generates every character, environment, and object reference once, and renders each shot with the right references attached. GPT Image 2.5 Flare is now the default image model for all of those steps, and the prompting rules above are how the storyboard writes its shots. If you make long-form video with consistent characters, how character consistency works across a whole video is the companion to this post, and you can start a project to see the model in that pipeline.

Share