Tool
Which model should you use?
Move the dials for what your work actually needs. Every model is rescored as you move them, against the same benchmark numbers, and the ranking reorders live.
Start from a job
What matters, and how much
Where it runs
Ranking
How the score works, and where the numbers come from
- Three sources, and a model is listed only if all three scored it. LiveBench 2026-06-25 gives seven category scores and cost per successful task. Artificial Analysis gives an intelligence index, cost, output speed and time to first token. Agent Arena gives four session signals from 1.4 million real sessions. That rule is the whole relevance filter: it keeps the list to models people actually choose between, and it means every row is measured the same way.
- Each dial is a weight. Turn one up and it takes a bigger share. The shares shown next to the dials add to 100. Only open weights is a filter, because a model either ships its weights or it does not.
- Scores are percentiles across the models that published that number, not the raw benchmark values, so a 90 means near the top of this field rather than 90%.
- One row per model, at its best published setting. The leaderboards disagree about which reasoning effort to publish: LiveBench scored Claude Fable 5 at max effort, Artificial Analysis scored its default. Nothing is ever copied from one setting to another; the row simply carries what each source actually measured.
- None of the three composites is imported. Arena's Net Improvement column is the number its own rank is built from, so copying it would replay their ranking instead of building one. Its Tool Hallucination column is left out too: one value fills fifteen of 44 rows and all ten places of its own top-ten list, which is a placeholder rather than a measurement.
- Arena prints its signals with the minus signs stripped. Read as printed, the worst-ranked model looks like the best at recovering from a failed command. Each signal's own top-ten list gives a threshold that recovers the sign, and that is done in
data/arena-agent.jsonbefore anything is scored. - Cost is per finished task, not per token. A cheap model that needs three attempts is not cheap.
Nothing matches. Turn off the open-weights requirement, or turn up a second dial.