← All articlesResearch

LMArena text-to-image leaderboard: 80 models, 6.5 million votes

Megan BrooksSeptember 30, 20267 min read
LMArena text-to-image leaderboard: 80 models, 6.5 million votes

The LMArena text-to-image leaderboard as of September 2026: GPT Image 2.5 on top, where open-weight models land, and what Elo can't tell you.

My notes on the text-to-image board at arena.ai, the site that used to be LMArena. The original lives at arena.ai/leaderboard/text-to-image. It's a live table and new votes land every day, so if you're reading this a month from now, check the current numbers there.

We were reshuffling the default set of models Flami suggests for product images, and I wanted an outside reference. Not our own taste. A big pile of strangers' eyeballs. So I went to the LMArena text-to-image leaderboard, which is still the most quoted ranking of image models, even though the old lmarena.ai address now just redirects to arena.ai. How it works is simple enough. Someone types a prompt, gets two images from two unnamed models, picks the one they like better, and all those wins and losses roll up into an Elo-style score. The September 24, 2026 update lists 80 models and 6,478,738 votes. Below is what I wrote down, with my own comments.

OpenAI now holds the first three spots

Back in June, GPT Image 2 was alone at the top with a lead of more than a hundred points. It's still up there. It just got pushed down to third by its own younger siblings.

  • 1st, gpt-image-2.5-sunburst (OpenAI), 1424, marked preliminary;
  • 2nd, gpt-image-2.5-flare (OpenAI), 1401, also preliminary;
  • 3rd, gpt-image-2 in medium quality (OpenAI), 1383;
  • 4th, mai-image-2.6 (Microsoft AI), 1335;
  • 5th, reve-2.1 (Reve), 1302.

OpenAI shipped Images 2.5 on September 8 (OpenAI's announcement), and developers got two API versions. Flare is the default, while Sunburst is slower and aimed at detailed work where you need tight editing control (9to5Mac has a good rundown). I'd hold off on crowning either one. Both have only about 10,000 votes each and a confidence interval of plus or minus 8 points, so the order between them could still shift. GPT Image 2 is a different story, with 88,744 votes and an interval of 4 points, so its 1383 is about as settled as numbers get on this board.

What really jumps out is the gap between OpenAI and everyone else. From Sunburst down to Microsoft's MAI-Image-2.6 it's 89 points. Google's best entry, Nano Banana 2 in its web-search mode, sits ninth at 1261. I broke down how GPT Image 2 built a lead like this across two different arenas in that piece. For what it's actually like on real jobs, there's the GPT Image 2 review.

The middle of the table, where the working models live

Zero points. That's the distance between seedream-5.0-pro and qwen-image-3.0-pro, tied at 1256 in 10th and 11th. Past the top ten, neighbors are usually a few points apart, and the confidence intervals there run from 3 to 6 points, so in practice those rows are interchangeable. What decides it is price, mostly, and whether the model even understands what your product is.

A few rows I care about. grok-imagine-image is 23rd with 1170 and 262,132 votes. In June it had more votes than anything else on the board. Now the record belongs to the original Nano Banana with 878,451, which probably just means it's been sitting there the longest. qwen-image-2.0-pro is 19th at 1192.

Seedream is the interesting one. The family split in two. The new seedream-5.0-pro jumped to 10th, while seedream-4.5 is down at 35th (1147) and seedream-5.0-lite at 38th (1138). Lighter versions always take a beating in blind tests, I think because voters see one frame and never the price tag. The whole lineup, side by side, is in the Seedream 5 review if you want the details; the model itself lives on the Seedream 5 page.

Where do open-weight models actually rank?

Depends on what you count as open, honestly. If you only count permissive licenses like Apache 2.0 or MIT, the best one is qwen-image-2512 in 42nd place with 1125. Right behind it is hidream-o1-image under MIT, 46th at 1116. Then there's a long stretch of closed models before z-image-turbo, also Apache 2.0, in 57th with 1084.

Loosen the definition to "weights you can download" and it looks a lot better. qwen-image-2.1 is 17th at 1228, though it's preliminary with just 4,445 votes and a research license. Ideogram 4.0 in quality mode is 18th at 1205. Ideogram published its weights on June 3, 2026 (their launch post), and it's the most voted open model this high up. Nvidia's Cosmos3 is 25th. Each one has its own license, and I'd read it twice before using any of them for client work.

wan2.7-image-pro is 53rd at 1103. It's listed as proprietary, not open, but I'm mentioning it because we keep it around for product shots, see the Wan 2.7 Image page or the Wan 2.7 Image review if you want to try it.

Fifty-seventh sounds bad for Z-Image Turbo. Thing is, the arena measures how nice one frame looks and ignores what a thousand of those frames cost you. Why Turbo stays our main draft engine despite that score is a longer story, and I got into it in the Z-Image write-up; you can also just try it yourself on the Z-Image page.

What the leaderboard won't tell you if you sell online

People voting on the arena are just picking the prettier picture. Nobody there is asking whether it'll actually sell your product. Nobody's checking whether the small print on your packaging stays readable, or whether the product in the shot actually matches your reference photo, and getting 300 SKUs done within budget isn't on that ballot either. None of that is on the ballot. I use the overall rank as a first filter and nothing more, then run the actual product through whichever two or three models look decent on paper. Maybe I'm overcautious here, but every model from this table still goes through our own product tests before we hand it to users.

If you want a quick sanity check of your own, you don't need 6 million votes. Take five of your real product photos, run the same prompt through the three or four models that look good on paper, and put the results next to each other without the names. A friend picking blind will tell you more about your catalog than any public board. It won't hold up as statistics, but at least you're judging your own catalog instead of a stranger's.

Everything from this snapshot that passed those tests is on Flami, and new accounts get free starter credits to try them.

Quick answers

Which AI image model is number one on arena.ai right now?

As of the September 24, 2026 update, it's gpt-image-2.5-sunburst from OpenAI with 1424. It's still marked preliminary. Flare is second with 1401, and GPT Image 2 is third with 1383.

How is arena.ai different from Artificial Analysis?

Same basic idea, blind pairwise votes turned into Elo. The voters differ, though, and so do the prompts. Positions don't line up, and the scores sit on different scales. I check both arena.ai and the Artificial Analysis leaderboard. For what it's worth, Sunburst is on top at Artificial Analysis too.

Are there open-weight models near the top?

Yes, if you accept custom licenses. qwen-image-2.1 is 17th and Ideogram 4.0 is 18th, both with their own licenses. The best Apache 2.0 model, qwen-image-2512, is 42nd.

What happened to LMArena?

It moved from lmarena.ai to arena.ai, and old links redirect to the same text-to-image board. The vote history seems to have come along too, since older models like the original Nano Banana still show hundreds of thousands of votes.

How is the score calculated?

A voter sees two images for the same prompt with no labels saying which model made which. Ratings get recalculated chess style from wins and losses. As votes pile up, the confidence interval shrinks. The current leader sits at plus or minus 8 points on roughly 11,000 votes. GPT Image 2 has eight times as many votes and sits at plus or minus 4.

Updated September 30, 2026, based on the September 24, 2026 leaderboard refresh.

Sources

  1. arena.ai, Text-to-Image Leaderboard
  2. Artificial Analysis, Text-to-Image Leaderboard
  3. OpenAI, Introducing ChatGPT Images 2.5
  4. 9to5Mac, OpenAI releases ChatGPT Images 2.5
  5. Ideogram, Ideogram 4.0

About the author

Megan Brooks

Reviewer at Flami

Read next

Get 15 credits for free

Use them to generate images and videos:

≈ 11 × Nano Banana 2, ≈ 12 × GPT Image 2, ≈ 1 × Seedance 2

Get free credits