- OpenAI released ChatGPT Images 2.5 on 8 September 2026 and put two models behind it in the API, gpt-image-2.5-flare and gpt-image-2.5-sunburst.
- The line in the release notes that matters to a shop is not sharper rendering. It is better preservation of subjects from reference photos, which is the whole ballgame for a catalogue.
- Templates now ship for formats including Poster and Merch, which tells you plainly which customers OpenAI is chasing.
- Generation latency is down by up to 50 percent against Images 2.0. That changes batch work far more than it changes a single hero shot.
- Published API rates are 30 dollars per million image output tokens and 8 dollars per million image input tokens, with a cached input rate of 2 dollars.
- None of this moves your disclosure duty. Marking the output stays with the tool provider under the AI Act, and your exposure stays where it always was, in whether the picture tells the truth.
A small shop does not have an image problem. It has a specific image problem, and anyone who has photographed their own stock on a kitchen table knows the shape of it. You need eleven pictures of the same navy tote from angles that do not exist in your flat, on a background you do not own, with light you cannot make in November. The generated image was supposed to solve that eighteen months ago and mostly did not, because the machine kept returning a beautiful navy tote that was not your navy tote.
That is the frame to read the 8 September 2026 release through. OpenAI shipped ChatGPT Images 2.5, and the coverage led with sharper detail and richer texture, which is the least interesting thing about it if you sell physical goods.
What actually shipped on 8 September 2026?
The rollout went to ChatGPT, ChatGPT Work and Codex users across every tier, on desktop, on mobile and on the web, according to 9to5Mac's write-up of the release. There is no waiting list and no separate paid add-on to find. If you already have an account, the model behind the image button changed under you.
In the API there are now two model identifiers rather than one. Simon Willison's notes on the release record them as gpt-image-2.5-sunburst, recommended where editing precision matters most, and gpt-image-2.5-flare, aimed at fast everyday generation. Willison also updated his own command line tool the same day to pass reference images with an -i flag, which is a small detail that tells you where the interesting work now sits: not in the prompt, in the photo you feed alongside it.
The feature list beyond the model itself is short. A drawing surface called Sketch lets you scribble a rough composition and hand it over as a visual reference. Templates give you a starting point for common formats, and the two named in Unite.AI's coverage of the two new API models are Poster and Merch. You can drop a comment onto a spot in a generated image to steer an edit at that spot rather than describing it in words. Prompts can be shared. OpenAI puts current volume at more than 3 billion images a week across ChatGPT Images and the API models.
Poster and Merch are not neutral choices of demo. They are the two jobs a person with something to sell does most often, and putting them behind a template button is a statement about who this release is for.
Why does reference fidelity matter more than image quality?
Because the failure that costs you money is not an ugly picture. It is a picture of the wrong object. A generated shot that looks like a magazine spread but shows a buckle you do not ship, a weave one grade finer than your fabric or a logo two centimetres off centre is worse than a phone snap on a bedsheet, and it is worse in a way you only discover through returns.
Every version of this technology has been good at plausibility and bad at identity. Ask for a leather satchel and you get a leather satchel. Ask for this leather satchel and, until recently, you got a cousin of it. The release notes for 2.5 claim better preservation of subjects from reference photos and more reliable instruction following across multiple turns of editing. Those two claims, if they hold on your stock, are the difference between a toy and a tool.
Multi-turn reliability deserves its own sentence. Catalogue work is never one prompt. It is a prompt, then a correction, then a second correction, and the old failure mode was that the third instruction quietly undid the first. If the model now holds an edit while applying the next one, the workflow stops being a slot machine.
Which catalogue jobs can a generated image actually do?
Not all of them, and the split is not about difficulty. It is about which attribute a customer can send the parcel back over. A background is not returnable. A colourway is. Here is the division we would apply to a real catalogue, built from what the model is now claimed to hold and what it is still free to invent.
| Catalogue job | What the model has to hold | What a mistake costs | Verdict |
|---|---|---|---|
| Background swap behind a real packshot | The silhouette only | Little. The product is your photo | Safe to automate |
| Lifestyle scene built around your product | Shape, colour, logo position | A return if the item drifts | Usable with a check on each output |
| Poster or merch mockup for an ad | Brand assets and rough proportion | Little. Nobody buys the poster | Safe to automate |
| A colourway you stock but never shot | An exact colour you cannot supply as reference | High. Colour is the top return reason in apparel | Photograph it |
| A new angle of an item shot once | Full three dimensional geometry | High. The model guesses the unseen side | Photograph it |
| Scale or size demonstration | Proportion against a known object | High, and it reads as deception | Photograph it |
| Packaging or label close-up | Legible small text | High. Wrong text on a label is a compliance issue | Photograph it |
The pattern in that table is worth stating out loud, because it is the rule you can carry into the next release without rereading anything. A generated image is safe where it adds context and unsafe where it supplies evidence. Context is the room the chair sits in. Evidence is the colour of the chair. Our longer piece on what marketplaces require from an AI generated product photo works the same divide from the policy side and lands in the same place.
What does it cost to run at catalogue scale?
The published API rates, as recorded in the Unite.AI breakdown, are 30 dollars per million image output tokens, 8 dollars per million image input tokens with a cached rate of 2 dollars, and 5 dollars per million text input tokens with a cached rate of 1.25 dollars. We are deliberately not converting those into a price per image here, because the number of output tokens depends on the size you request and inventing a per-image figure would be exactly the sort of thing this blog gets to be wrong about in public.
What you can reason about without a calculator is the shape of the bill. Image output dominates it, at nearly four times the image input rate. So a workflow that sends one reference photo and asks for six variants is cheap on the way in and expensive on the way out, and the lever that matters is generating fewer, better candidates rather than spraying thirty and picking one. Caching cuts the input side by three quarters, which rewards a stable house prompt reused across a catalogue instead of a hand written prompt per product.
Latency is the other half of the economics and it is easy to underrate. A 50 percent cut against Images 2.0, with the faster model claimed at two to four times the speed of its predecessor, does very little for one hero image. For a run of 400 SKUs it is the difference between an overnight job and an afternoon, which is the difference between doing it and not doing it. Speed is a scheduling feature, not a quality feature.
Does a better model change what you have to disclose?
No, and the reason is worth understanding once so you stop worrying about it. Under Article 50 of the EU AI Act, which has applied since 2 August 2026, paragraph 2 puts the marking duty on the provider of the system. It is OpenAI that has to ensure output is marked in a machine readable format and detectable as artificially generated. That obligation does not travel to you when you press the button.
Paragraph 4 is the one people misread. It obliges a deployer to disclose when the content is a deep fake, and it carries a softer regime for artistic or fictional work. A studio shot of your own kettle, generated from a photograph of your own kettle, is not a deep fake by any reading of that paragraph. The duty you are actually exposed to is older and more boring: the picture must not misrepresent the goods, which is ordinary consumer protection law and predates every model in this article.
Where the real operational risk sits is metadata survival, and we have written that up separately in the piece on what a content credential actually proves about a file. The provider marks the file. Your resize step strips the mark. Nobody decided anything and the label is gone.
There is a wrinkle in paragraph 2 that is worth holding on to, because it may matter more as this workflow spreads. The marking duty carries an exemption for systems performing an assistive function for standard editing which does not substantially alter the input data or its meaning. A background swap around an untouched packshot is a reasonable candidate for that description. A fully generated scene is not. We are not going to tell you where the line falls, because nobody has tested it yet, but if you are keeping records of which images were generated and which were retouched, that distinction is the one worth recording.
Who inside a one person business does this help?
The person who cannot describe a photograph in words. That is a real and underrated barrier, and it is the barrier Sketch is aimed at. You draw a rough box where the product goes and a smear where the light comes from, invoke it with an at sign, and hand the machine a composition instead of a paragraph trying to specify one. Anyone who has typed three quarter view, soft key light from the left and received something else entirely will recognise what that removes.
Comment placement does something similar at the other end. Rather than writing a fresh prompt that describes the whole image again in order to change one corner of it, you put a note on the corner. For a seller iterating on a single listing image that is the difference between four attempts and one, and it compounds with the multi turn reliability claim: the corrections now have somewhere specific to land.
Prompt sharing is the quiet one. If you work with a freelancer or a virtual assistant, a house prompt that produces your catalogue look is an asset, and until now it lived in somebody's notes app. Shared prompts plus the cached input rate point the same direction, which is towards one stable house prompt reused across every product rather than improvisation per item. That is also how you get consistency, which is the thing customers actually notice across a catalogue page.
OpenAI has also put the model into an Adobe Firefly integration, which matters mainly as a signal about distribution. The image models are arriving inside the tools people already open rather than as a destination of their own. For a merchant that means the relevant question stops being which generator to sign up for and starts being which of your existing subscriptions quietly gained one.
What should you test before you cancel the photoshoot?
Do not evaluate this on your easiest product. Everyone does, everyone gets a delightful result, and the delight does not survive contact with the catalogue. Run it on the three items most likely to break it, then decide.
- Pick your three worst subjects. Something reflective, something with fine printed text and something with an unusual silhouette. Those three cover most of what breaks.
- Give it one reference photo each, then generate ten variants. One reference, because that is the real constraint of a small shop. Ten variants, because you are measuring consistency rather than best case.
- Score drift on the returnable attributes only. Colour against the item in your hand. Texture. Proportion. Logo position. Legibility of any printed text. Ignore whether the picture is pretty.
- Publish the passing ones to a slice of the catalogue and leave the rest alone. A quarter of live traffic tells you more than any amount of side by side comparison in a browser tab.
- Read the return reasons, not the conversion rate. Conversion goes up when pictures get prettier. Returns tell you whether the picture was true. If returns for the wrong item or wrong colour move on the generated slice, stop.
That protocol takes about a fortnight of elapsed time and almost no work. It also gives you a defensible answer when a marketplace asks why your imagery changed, which is a question sellers do get asked.
What this release does not fix
Small text on labels and packaging is still the weakest part of every image model, and it fails in a specific way that is dangerous for commerce: it produces text that looks right at thumbnail size and is wrong when a customer zooms. Ingredients, wash symbols and certification marks all live in that danger zone.
Colour remains a supply chain problem rather than a model problem. Your screen, the model's notion of navy and the dye lot that arrived from your supplier are three different navies, and no amount of reference preservation reconciles them. If colour accuracy decides your returns, a physical shot under known light is still the only honest answer.
And a faster model does not touch the part of the job that takes the longest, which is deciding what a listing should say and show. Our notes on which parts of a marketplace listing to override by hand cover that ground. Neither does it help with the print side of things, where a generated mockup and a real garment famously diverge, as the walkthrough of how a print on demand file differs from the mockup sets out.
So what would we actually do with it?
Two things this month, if we ran a small catalogue. Move every background swap to the model and stop paying for that by the hour, because it is the job with the lowest downside and the model has been good enough at it for a year. Then run the fortnight test above on lifestyle scenes, which is the job with real upside and real risk, and let the return data decide rather than the demo.
What we would not do is treat a release note as evidence about our own products. Better subject preservation is a claim about an average across a training distribution. Your stock is not an average, and the only dataset that matters is the one sitting in your stockroom. If you are building the storefront these images will live on, our AI ecommerce store builder generates the product pages and you keep the code, which at least means the image pipeline is yours to inspect when metadata starts disappearing from it.
The honest summary of 8 September is that image generation moved one notch closer to being a catalogue tool rather than a mood board tool, and the notch it moved on was the right one. That is a genuinely useful release. It is not a reason to sell your camera.