← All articlesResearch

GPT Image 2 leaderboard check: it won two arenas, then OpenAI outranked it

Megan BrooksSeptember 30, 20266 min read
GPT Image 2 leaderboard check: it won two arenas, then OpenAI outranked it

GPT Image 2 on the Artificial Analysis and arena.ai leaderboards: how big its lead was, where it ranks in September 2026, and what the Elo gaps mean.

This is my read of two independent text-to-image leaderboards, Artificial Analysis and arena.ai. Both recalculate after every batch of votes, so treat the numbers below as a September 2026 snapshot and check the originals for today's figures.

88,744 votes. That's how many GPT Image 2 has collected on the arena.ai leaderboard by now, and its score has barely moved since it had half as many. Across both big arenas, GPT Image 2 sits third. Third sounds like a demotion, and in a way it is, but the two models above it are OpenAI's own. I do art direction at Flami, so choosing which model makes the hero image is literally my call, and I've been watching this one since spring. What caught my eye back then was rare: two arenas that have nothing to do with each other agreeing on the same winner. What's interesting now is how that lead shrank.

Why I bother with two arenas

A single arena can be wrong. Who shows up to vote matters, and so does the pool of prompts. Model pairing skews it too, more than people think. Any of that nudges the ranking. If two boards with different methods put the same model on top, I take that far more seriously than either one alone.

Both use blind pairwise voting. If you want the mechanics, I walked through them in my notes on the arena.ai image board. And the scale is serious. arena.ai now has 80 models and almost 6.5 million votes in its text-to-image category, as of its September 24, 2026 update. Random luck doesn't survive numbers like that.

Artificial Analysis: from a 74-point lead to 17

In June the picture on Artificial Analysis was lopsided. GPT Image 2 (high) led by 74 Elo over the runner-up, which was more than the whole distance from second place to fifth.

Today it reads differently. The top of the board:

  • 1st, GPT Image 2.5 Sunburst (max), 1197;
  • 2nd, GPT Image 2.5 Flare (max), 1190;
  • 3rd, GPT Image 2 (high), 1172 on 15,346 comparisons;
  • 4th, Grok Imagine Image 2.0 from xAI, 1155;
  • 5th, MAI-Image-2.6 from Microsoft AI, 1150.

One caveat before anyone compares this with older screenshots. The board now runs on a revised methodology labeled v2.0, and the scores sit on a different scale. The 1340 GPT Image 2 posted in June and the 1172 it has now aren't the same ruler, so don't read that as a 170-point collapse. What you can compare is the gap. Against the best non-OpenAI model it's down to 17 points, with a 95% interval of plus or minus 9 on each side. In practice that's 52 wins out of 100, not the runaway lead the June numbers showed.

Further down, Nano Banana 2 is sixth at 1125 and GPT Image 1.5 (high) eighth at 1107. Among open-weight entries, Qwen-Image-2.1 is highest, 18th with 1034. In June that title belonged to Cosmos3-Super, for what it's worth.

arena.ai: the "preliminary" tag came off, the score held

This is the part I find reassuring. In June, gpt-image-2 (medium) sat at 1385 with 45,100 votes and a Preliminary flag, which means the rating could still slide as votes came in. I guessed at the time it might lose 30 points or so. It didn't. On the September 24 update it's at 1383 plus or minus 4 with 88,744 votes, flag gone.

It lost first place anyway. On September 8 OpenAI released Images 2.5 for ChatGPT (here's the launch post). Its two API versions went straight in above their older sibling. gpt-image-2.5-sunburst has 1424 and gpt-image-2.5-flare 1401, both still preliminary with roughly 10,000 votes each.

Where it gets interesting is the gap to everyone outside OpenAI. Back in June the runner-up was reve-2.0 at 1273, more than a hundred points behind, and I couldn't remember a lead that size on an image arena. Now the closest challenger is Microsoft's mai-image-2.6 at 1335. That's 48 points back, while reve-2.0 has slipped to eighth. OpenAI is still ahead, just by a margin Microsoft could close with one solid release.

What does an Elo gap actually mean?

Elo comes from chess, and the arithmetic behind it is friendlier than it looks. A 100-point gap means the higher-rated model should win about 64 of every 100 head-to-head votes. At 48 points, GPT Image 2's current margin over Microsoft on arena.ai, it's closer to 57 out of 100. At 17 points, its edge on Artificial Analysis, you're down to about 52.

So yes, GPT Image 2 still beats nearly everyone. It just beats them more narrowly than in spring, and voters would struggle to tell the top few apart in a lot of pairs.

And the models lower down

Worth a glance at the bottom half of the table too. On arena.ai Seedream 5.0 Lite is 38th with 1138. In June it was 28th, and the score barely changed, so newer models simply stacked up above it. If you want to test it on product shots, Seedream 5 is available in Flami, and the whole family is compared in a separate piece.

Z-Image Turbo is 57th with 1084, down from 46th. Same blind spot as always, the arena doesn't credit it for running on your own hardware. I go through why that still matters in the LMArena breakdown. We went into that trade-off in our Z-Image research piece, and you can try it on Flami directly.

A low rank there just means fewer strangers happened to like that one frame, not that it's wrong for your task.

What I actually do with this

For a single hero image where quality is the only thing that matters, my first pick is still OpenAI's model. It took both arenas in April. TechCrunch's launch coverage singled out how well it renders text, and that's exactly the detail that matters most on packaging shots. I'm not sure yet whether 2.5 is worth switching to for everyday work. The preliminary scores say yes, but I'd probably wait until they settle.

Easiest way to judge for yourself is to upload one of your product photos to Flami and run the same brief through two or three of these models side by side.

Quick questions

  • Where does GPT Image 2 rank right now? Third on both boards, behind OpenAI's own Images 2.5 pair. On arena.ai it scores 1383 on 88,744 votes. On Artificial Analysis it's 1172 on 15,346 comparisons.
  • Why did its Artificial Analysis score drop from 1340 to 1172? The board switched to a revised v2.0 methodology with a different scale. Compare gaps between models, not raw scores across versions.
  • What did the Preliminary tag on arena.ai mean? It flags a model that hasn't collected enough votes yet. GPT Image 2 lost the tag over the summer, and its score moved by only 2 points.
  • Which open-weight model ranks highest on Artificial Analysis? Qwen-Image-2.1, 18th with 1034. On arena.ai the picture depends on how strictly you define open, and I'd check each license before client work.

Snapshot taken September 30, 2026, from the arena.ai update of September 24, 2026 and the live Artificial Analysis board.

Sources

  1. TechCrunch, ChatGPT's new Images 2.0 model is surprisingly good at generating text
  2. OpenAI, Introducing ChatGPT Images 2.0
  3. Artificial Analysis, image arena rankings (text to image)
  4. OpenAI blog, ChatGPT Images 2.5 launch
  5. arena.ai, Image Arena, text-to-image board

About the author

Megan Brooks

Reviewer at Flami

Read next

Get 15 credits for free

Use them to generate images and videos:

≈ 11 × Nano Banana 2, ≈ 12 × GPT Image 2, ≈ 1 × Seedance 2

Get free credits