OpenAI 推出的 GPT-Images-2.5
OpenAI 发布了 ChatGPT Images 2.5 下的两款全新图像模型,宣称细节更锐利、编辑更精准、生成速度更快。绘图工具与可分享提示词等新功能,旨在拓展创作流程的边界。
OpenAI 表示,目前用户每周通过 ChatGPT Images 以及 API 中的 GPT-Image 模型生成的图片已超过 30 亿张。借助 ChatGPT Images 2.5,该公司正在推出重构后的生成能力,宣称可带来更自然的光影与更细腻的质感。
这两款模型旨在更好地保留参考照片中的主体特征,并在多轮编辑中更可靠地遵循指令。据 OpenAI 称,与 Images 2.0 相比,速度更快的版本可将图像生成延迟最高降低 50%。
两款新模型,各司其职
面向开发者,OpenAI 将这两款模型都带到了 API 中。GPT-Image-2.5 Flare 是大多数场景下的默认选择,图像质量高于 GPT-Image-2,同时延迟降低 50%。更强的 GPT-Image-2.5 Sunburst 则面向要求更高的视觉任务,对编辑过程有更精细的控制,但需要更长的生成时间才能交付结果。
两款模型采用相同的 token 费率:每百万图像输入 token 8 美元,每百万输出 token 30 美元。但由于 token 用量因模型和质量档位而异,相同的费率并不意味着单张图像的成本相同。Images 2.5 新增了“xhigh”和“max”两个质量档位,突破了此前“high”档的上限。
一张 1024x1024 的图像在“low”档下成本约为 0.006 美元,与前代持平;而“high”档约为 0.053 美元;新增的“max”档则约为 0.21 美元,对应约 7,024 个输出 token。这使得 Images 2.5 的“max”档与 GPT-Image-2 的“high”档价格相当。与前代不同,Images 2.5 目前没有更便宜的批量费率。不过在早期测试中,尽管 token 价格相同,Sunburst 的单张图像成本通常高于 Flare,这可能是由于其更长的推理过程所致。此外,与上次不同,OpenAI 没有给出平均单张图像价格的数据。
目前尚不清楚 OpenAI 如何在两款模型之间为 ChatGPT 用户分配流量。无论是公告还是文档,都没有说明 ChatGPT 何时会调用更快的 Flare 或更精准的 Sunburst。在 API 中,用户可以显式选择模型,但 ChatGPT 界面目前还没有提供这样的控制选项。
在我们的测试中,这条分界线目前似乎主要存在于 Chat 和 Work 之间。在 Work 中,无论推理设置如何,我们的提示词确实只会改变所要求的内容——这些定向编辑正是本次模型改进的重点。而在 Chat 中,即使将推理设为 high,后续图像中仍有更多细节会不断发生变化。只有在“6 Pro”设置下,更强的模型有时才会发挥作用。
只改变你所要求内容的编辑
如前所述,新模型的重点是编辑。它只改动被要求的元素,而让图像的其余部分保持不变,即使面对更复杂的主题和背景也是如此。在较长的多轮对话中,之前的修改应保持一致,且图像质量不会在多次编辑步骤中下降。OpenAI 以房间改造为例展示了这一点。旧模型在每次编辑时都会改动其他细节,而新模型即使经过多次迭代也能保持稳定。
我们通过 ChatGPT Work 中的 GPT-6 Astra(Max)做了一次快速测试,看看效果如何。我们使用了下面的提示词,然后迭代式地调整了一个小细节(香蕉的颜色)和一个大改动(那只大猫)。
一张超写实的 DSLR 照片。前景中,一只猴子拿着一根粉红色的香蕉,坐在一只老虎身上。背景中,一匹马正骑在一名宇航员身上。宇航员在下方,就像一个活生生的“宇航服马鞍”,而那匹马明显在上方、处于掌控地位,是骑手。请做到 100% 毫无歧义:马是骑手,宇航员是被骑的一方,绝不是反过来。高分辨率、对焦清晰、光线写实。

图片:GPT-Images-2.5 / Astra(Max)
顺便一提,这可能是我们测试中 OpenAI 图像模型生成过的最好的一版“马骑宇航员”画面。以下是 Image 2.0(Thinking 变体)的表现。

相比之下,在 ChatGPT 的 Chat 模式中进行的测试里,每当我们调整香蕉的颜色,图像的其余部分也总会随之改变。只有在“6 Pro”模式下,图像才在一次运行中保持了前后一致,但在另一次运行中却又不行。随着本周功能的逐步推出,这种情况可能会发生变化。
Images 2.5 还旨在更好地处理复杂的视觉指令,为现实世界的信息提供更准确的内容,并支持透明背景和更复杂的版式要求。例如,在我们上一篇文章中,我们让 ChatGPT 把文章做成了 80 年代杂志的跨页排版。效果如下。

通过 Astra(Max)调用的 GPT-Images-2.5 也根据我们“猴子宇航员”的提示词,生成了包含示例图片的详细杂志页面,并提供了较弱和较强两种变体,同时即便在质量较差的设定下,也没有偏离原始指令。
该模型能按指令纠正我们故意植入的一个错误(第一页图片上方写着 GPT-Images-2.5 而非 2.0),显示出它在保留其余内容方面表现出色;同样,当我们简单要求“生成一个美式英语版本”时,它也能在翻译时做到这一点。
图片:GPT-Images-2.5 / Astra(Max)
作为对比,通过普通 ChatGPT Chat 得到的结果也值得一看,但它们表明,较弱的模型在翻译过程中也会改动其他细节。不过,在某个案例中,它反而更准确地还原了微软现任 CEO 的形象。
图片:GPT-Images-2.5 / ChatGPT Chat(中等与即时)
绘图、模板与可分享的提示词
在 ChatGPT 中,OpenAI 随新模型一同推出了多项新功能。其中最重要的是名为“Sketch”的功能。它让用户可以直接在 ChatGPT 中绘图,并将草图作为成品图像的视觉模板。用户通过“@Sketch”命令激活该功能,OpenAI 表示它非常适合制作图表、房间布局或海报。
模板是预设好的提示词,旨在让用户更轻松地开始制作海报、标志、信息图、缩略图、插画或广告等格式的内容。它们提供的是结构而非空白画布,并通过有针对性的追问来锁定你的需求。用户还可以直接在图像上添加评论,并分享自己使用的提示词,这样其他人就能用自己的照片和细节尝试同样的创意。作为示例,OpenAI 提到了一个目前正在病毒式传播的提示词,它可以生成 1980 年代风格的人像。
两款模型登顶 Arena 排行榜
在 Arena 的文生图排行榜上,新模型目前占据前两名。GPT-Image-2.5 Sunburst 以 1421 分领先,紧随其后的是 GPT-Image-2.5 Flare,得分 1399。两个分数都带有“初步”标签,并且基于仍然较低的投票数——分别约为 3100 票和 2900 票。排名第三的是前代产品 GPT-Image-2,得分为 1381,基于约 78,700 票。
紧随其后的是微软的 mai-image-2.6(1331 分)、SpaceXAI 的 grok-imagine-image-2.0(1315 分),以及来自 Reve、Meta、Google 和字节跳动的多个模型。由于 OpenAI 这两款新模型的评分仍是初步结果,随着更多投票涌入,它们的排名可能还会发生变化。
与 Google DeepMind 合作的水印技术
在来源标注方面,OpenAI 仍依赖 C2PA 行业标准,该标准通过嵌入元数据来实现追踪。为了让来源标注更具抗篡改能力,OpenAI 还在 ChatGPT、Codex 和 API 中通过 Google DeepMind 的 SynthID 添加了不可见水印。OpenAI 表示,来源标注没有单一的解决方案,因此采取了分层式的方法。
Images 2.5 现已面向全球所有 ChatGPT、ChatGPT Work 和 Codex 用户在桌面端、移动端和网页端推出,包括免费版用户。OpenAI 表示,等待时间更短,使用限额也更高。我们的测试表明,Chat 模式目前使用的是较弱的模型。
GPT-Images-2.5 prompted by OpenAI
OpenAI has released two new image models under ChatGPT Images 2.5, promising sharper details, more precise editing, and faster generation. New features like a drawing tool and shareable prompts aim to expand the creative process.
OpenAI says users now generate more than three billion images each week through ChatGPT Images and the GPT-Image models in the API. With ChatGPT Images 2.5, the company is rolling out a reworked generation that promises more natural lighting and finer textures.
The models are meant to preserve subjects from reference photos better and follow editing instructions more reliably across multiple rounds. According to OpenAI, the faster variant cuts image generation latency by up to 50 percent compared to Images 2.0.
Two new models built for different jobs
For developers, OpenAI brings both models to the API. GPT-Image-2.5 Flare is the default pick for most uses, with higher image quality than GPT-Image-2 at 50 percent lower latency. The stronger model, GPT-Image-2.5 Sunburst, targets more demanding visual work with tighter control over edits, and it needs longer generation times to deliver.
Both models use the same token rates. That's eight dollars per one million image input tokens and 30 dollars per one million output tokens. But since token use varies by model and quality tier, the same rates don't mean the same cost per image. New in Images 2.5 are the "xhigh" and "max" quality tiers, which go beyond the previous ceiling of "high."
A 1024x1024 image at the "low" tier costs about 0.006 dollars, same as the predecessor, while "high" runs about 0.053 dollars, and the new "max" tier lands at roughly 0.21 dollars with around 7,024 output tokens. That puts the "max" tier of Images 2.5 at the same price as the "high" tier of GPT-Image-2. Unlike the predecessor, Images 2.5 has no cheaper batch rate so far. In early tests, though, Sunburst usually costs more per image than Flare despite identical token prices, probably because of longer reasoning runs. And unlike last time, OpenAI gives no average price-per-image figure.
It's also unclear how OpenAI routes ChatGPT users between the two models. Neither the announcement nor the documentation says when ChatGPT reaches for the faster Flare or the more precise Sunburst. In the API you can pick the model explicitly. The ChatGPT interface offers no such control yet.
In our tests, the line currently seems to run mainly between Chat and Work. In Work, our prompts really do change only what's asked, no matter the reasoning settings. These targeted edits are the focus of the model improvements. In Chat, though, more details keep shifting in the follow-up images, even with reasoning set to high. Only at the "6 Pro" setting does the stronger model sometimes appear to kick in.
Editing that changes only what you ask
As noted, the new model's focus is editing. It's meant to change only the requested elements and leave the rest of an image untouched, even with more complex subjects and backgrounds. In longer conversations, earlier changes should stay consistent without image quality dropping over multiple editing steps. OpenAI shows this with a room redesign, for example. Where the older model changed other details with every edit, the new model stays stable even across several iterations.
A quick test through ChatGPT Work with GPT-6 Astra (Max) shows how well this works. We used the following prompt and then iteratively adjusted one small detail (banana color) and one large one (the big cat).
A hyper-realistic DSLR photo. A monkey holding a pink banana is sitting on a tiger in the foreground. In the background, a HORSE is RIDING AN ASTRONAUT. The astronaut is underneath, like a living “spacesuit horse saddle,” and the HORSE is clearly on top, in control, as the rider. Make it 100% unambiguous: the HORSE is the rider and the ASTRONAUT is being ridden, NOT the other way around. High resolution, sharp focus, realistic lighting.

Image: GPT-Images-2.5 / Astra (Max)
On a side note, this is probably the best version of a horse riding an astronaut that an OpenAI image model has produced in our tests. Here's how it looked with Image 2.0 (Thinking variant).

Tests in ChatGPT's Chat mode, by contrast, always changed the rest of the image when we adjusted the banana color. Only in "6 Pro" mode did the image stay consistent in one run, but not in another. This may change over the course of the week as the rollout continues.
Images 2.5 is also meant to handle complex visual instructions better, deliver more accurate content for real-world information, and work with transparent backgrounds and more demanding layouts. In our last article, for instance, we had ChatGPT turn the piece into an 80s magazine spread. It looked like this.

GPT-Images-2.5 via Astra (Max) also produced a detailed magazine with a sample image based on our monkey-astronaut prompt, in a weaker and a stronger variant, without losing sight of the original instruction for worse quality.
The model corrected an error we planted (GPT-Images-2.5 instead of 2.0 right above the images on the first page) on command, showing how well it preserves the rest of the content, and it did the same for a translation on the simple request "Create a US English version."
Image: GPT-Images-2.5 / Astra (Max)
For comparison, the results through plain ChatGPT Chat are also worth a look, but they show that the weaker model changes other details during translation too. In one case, though, it nailed the current Microsoft CEO better.
Image: GPT-Images-2.5 / ChatGPT Chat (Medium and Instant)
Drawing, templates, and shareable prompts
In ChatGPT, OpenAI is adding several new features alongside the new models. The most important one is called "Sketch." It lets users draw directly in ChatGPT and use the sketch as a visual template for the finished image. You activate it with the "@Sketch" command, and OpenAI says it works well for diagrams, room layouts, or posters.
Templates are ready-made prompts meant to make it easier to get started with formats like posters, logos, infographics, thumbnails, illustrations, or ads. They give you a structure instead of a blank canvas and pin down your requirements through targeted follow-up questions. Users can also place comments directly on images and share the prompts they used, so others can try the same idea with their own photos and details. As an example, OpenAI points to a currently viral prompt that generates portraits in 1980s style.
Both models top the Arena ranking
In Arena's text-to-image leaderboard, the new models hold the top two spots for now. GPT-Image-2.5 Sunburst leads with a score of 1421, followed by GPT-Image-2.5 Flare at 1399. Both scores carry a "Preliminary" label and rest on still-low vote counts of around 3,100 and 2,900. In third place is the predecessor GPT-Image-2 with 1381 points from about 78,700 votes.
Behind them come Microsoft's mai-image-2.6 (1331), SpaceXAI's grok-imagine-image-2.0 (1315), and several models from Reve, Meta, Google, and Bytedance. Since the ratings for the two new OpenAI models are preliminary, their position could still shift as more votes come in.
Watermarking in partnership with Google DeepMind
For provenance labeling, OpenAI still relies on the C2PA industry standard, which embeds metadata for tracking. To make provenance more resistant, OpenAI also adds an invisible watermark via Google DeepMind's SynthID, in ChatGPT, Codex, and the API. OpenAI says there's no single solution for provenance labeling, which is why it takes a layered approach.
Images 2.5 is available now worldwide for all ChatGPT, ChatGPT Work, and Codex users across desktop, mobile, and web, including the free version. OpenAI says shorter wait times and higher usage limits apply. Our tests suggest the Chat mode currently uses the weaker model.






