Model benchmarks

Compare

Key models, from code to prose to generated screens. Built from the benchmarks that answer each question most honestly, with a link to the source for every number.

Who leads what

As of August 28, 2026

Overall

Which model leads overall, when coding, agents, knowledge and science count together?

Leaders

  1. 1Claude Opus 563
  2. 2Claude Fable 562
  3. 3Grok 4.661

Scale 0 to 100

Any single benchmark overfits something, so the index blends nine hard tests into one number. Start here if you want an allrounder. For one specific strength, read the other dimensions.

Based on

Nine underlying evaluations across four pillars: agents, coding, general capability and scientific reasoning. Each pillar carries a quarter of the weight.

The one composite that weighs agentic work on par with knowledge. The runner also publishes price per task and output speed, which often decides more than the score.

Agents and code

Which model carries a real task through terminals and repositories without me reworking every step?

Leaders

  1. 1Claude Opus 512.99%
  2. 2Claude Fable 511.7%
  3. 3GPT-5.6 Sol10.24%

Scale 0% to 16%

A high arena rank comes from observed runs, a high SWE-bench value from passing test suites. Together they bracket reality well.

Based on

People compare blind agent runs. The score is the share of tasks solved end to end.

Measures long tool using runs rather than answer sprints, which is closest to how coding agents actually work.

SWE-bench VerifiedSWE-bench Team

500 real GitHub issues from twelve Python repositories. A fix only counts when the test suite passes.

Long the standard for bug fixing on real code, by now close to saturated: frontier models sit around 96 percent. It is the floor check here, while the arena shows the differences.

Web and UI

Which model builds screens that look polished and professional, survive many variations and refuse the template look?

Leaders

  1. 1Kimi K31,372
  2. 2Muse Spark 1.21,335
  3. 3Claude Opus 51,332

Scale 1,200 to 1,450

The two boards crown different winners. Design Arena picks Kimi K3, the larger LMArena WebDev ranks Claude Opus 5 first at 1691 ahead of Kimi K3 Max. Where pure design decides, weigh Design Arena higher. For full web tasks, trust the arena.

Based on

People compare two rendered outputs of the same prompt: websites, UI components, data visualisations and 3D, each a single file.

So far the only benchmark that judges fully rendered design. It shows whether a result looks polished or like generic AI output.

Elo ranking from blind comparisons of real web development prompts, asked by real people.

The largest voter base among web benchmarks, and it crowns a different winner than Design Arena. That contradiction is why both belong here.

Writing

Which model writes prose with voice and rhythm instead of filler phrases, bullet lists and paragraph templates?

Leaders

  1. 1Claude Fable 51,505
  2. 2Claude Opus 4.6 High1,500
  3. 3Gemini 3.7 Flash High1,494

Scale 1,350 to 1,600

The top three sit within eleven Elo points, tighter than any other dimension. The rank here is a tendency, not a verdict. LiveBench complements the taste votes with fresh, controlled tasks.

Based on

The writing column of the text leaderboard: blind comparisons of creative and factual writing from real users.

Language quality where people actually vote, which makes it hard to game.

Continuously refreshed tasks, so answer patterns cannot leak into training. Includes a language and writing block.

The counterpart to taste voting: controlled criteria, immune to stale training data.

Languages

Which model writes German and Korean like a native speaker, not like a translation from English?

Leaders

  1. 1Gemini 3 Pro1,521
  2. 2Muse Spark 1.21,510
  3. 3Claude Opus 4.6 High1,509

Scale 1,400 to 1,600

Both boards measure writing in the language itself, not translation. They crown different winners: Gemini 3 Pro leads in German, Claude Fable 5 in Korean at 1501 ahead of Claude Opus 5. This site ships in both languages, so both boards count.

Based on

The German column of the text leaderboard: blind comparisons of prompts that users asked in German.

German is this site's first language. The board is the only one that shows whom German users prefer in blind comparisons.

The Korean column of the same leaderboard, with far fewer votes than the English columns.

Korean is this site's third language and the harder test: less training data, fewer votes, wider error bars.

Images

Which model generates images that work on assignment and go beyond stock photography?

Leaders

  1. 1GPT Image 21,382
  2. 2Mai Image 2.6 Preview1,331
  3. 3Grok Imagine 2.01,316

Scale 1,200 to 1,450

Aesthetics cannot be unit tested, so many people decide here. For campaigns and logos, Design Arena and LMArena keep separate SVG and logo categories.

Based on

Blind pairwise battles: two images, one prompt, one vote.

The only large sample measuring image quality against what people actually ask for, beyond lab marketing.

Vision

Which model reads screenshots, charts and long documents reliably enough for everyday browser work?

Leaders

  1. 1Claude Fable 51,313
  2. 2Claude Opus 4.7 High1,301
  3. 3Qwen 3.8 Max1,300

Scale 1,200 to 1,400

Vision does not replace your own eyes for pixel level design review. It does reliably sort which screenshots deserve a look.

Based on

Blind comparisons on image tasks: charts, screens and documents.

The everyday test for every screenshot that passes through a browser or tool.

Document understanding across long, messy files, the kind real paperwork leaves behind.

Complements screenshots with the other half of everyday work: PDFs, forms and contracts.

Quirks

Noted from daily work, not measured. Only what repeated often enough to be a pattern makes the list, with an arrow to the dimension it qualifies.

Claude Opus 5

Anthropic

Private vocabularyWriting

Wants to sound smart and gets hard to read: it invents terms, compresses ideas into abstractions and leans on words like load bearing and gates. The community calls these Claudisms.

GPT-5.6 Sol

OpenAI

Writes too much code and covers every line with tests until the actual diff drowns in test noise. Launch reviewers already called it overengineering.

Gemini 3 Pro

Google

Plans, then stumblesAgents and code

Reasons beautifully and stumbles in execution: the plan reads better than the patch that comes out of it.

Kimi K3

Moonshot AI

Needs an anchorWeb and UI

Strong once a document sits in context, flat without one. Chinese is first rate, English merely solid.

How the selection came about

The set holds what actually decides anything in my own work. The Intelligence Index is the compass for breadth, the arenas are the taste test, SWE-bench and LiveBench keep taste and stale task pools in check.

Elo tables weigh human preference: a top rank does not mean correct, it means frequently preferred. Models appear in their best documented effort level. Composite indexes bundle several tests and link each one.

Sources

The numbers come from these leaderboards, pulled on the reference date and maintained by hand.