In more detail
Arena Elo borrows the chess rating system to rank AI models by public blind taste test. People ask a real question, see answers from two unnamed models side by side, and vote for the better one. Win votes and the rating climbs; lose and it falls — across millions of matchups.
It matters because it’s hard to game. A model can cram for a fixed benchmark, but it can’t study for whatever real people happen to ask next — and voters don’t know whose answer they’re judging. That makes arena ratings the closest thing AI has to an honest public leaderboard.
Goes with
