Arena Leaderboard

Historical data from the Arena leaderboard, aggregated and visualized here from lmarena-ai/leaderboard-dataset on Hugging Face, licensed CC BY 4.0 — data current through . For live, up-to-the-minute rankings see the official Arena leaderboard ↗.

Category Leadership Timeline

Which lab held the #1 spot in each category leaderboard, over time, across all arenas (text, vision, WebDev, image/video generation, agent, and more). The eight labs with the most #1 finishes get their own color; every other lab (brief leads, early open-source models, etc.) is grouped under "Other" — hover a segment to see exactly who and which model. Click a group name to collapse or expand its subcategories.

One segment per leadership run (rank #1 model's lab) within a category. The current leader in each row is drawn slightly wider than its true date range so it stays clickable. Data through .

Rating History by Lab

Each line is one lab, tracking its best-performing model's rating over time (the max across all of that lab's models on the leaderboard that day). Pick a category, toggle up to 8 labs, and hover the chart — the tooltip names which model was in the lead.

Not directly comparable across categories. Data through .

Days Behind the Frontier

How far the leaderboard has to be rewound before each lab's best model from today would sit on top of it. The lab currently leading a category is 0 days behind; a lab at 180 means its best model would have led that category about six months ago. Hover a cell for the date behind the number.

Roll the field back to a past date — would this lab's best current model have been #1 then? The number is the most recent date where it would, which is simply the time since the first model that beats it was released. Which model beats which is judged at the most recent measurement of the pair, since a model's rating keeps moving as votes accumulate; models since delisted are placed on today's scale through the models they were last measured alongside. Where a lab actually held #1 more recently than that reconstruction implies, the observed finish is used instead. › N means the lab trails by more than the leaderboard's whole history for that category; “—” means the lab has no model there today.