jargon

Applied AI·Evaluation and reliability

the model at the top of the arena lost to the second one on your task, because humans voting on style is not your workload.

Leaderboards and Elo

Draft summary, pending review

Rankings such as LMArena, where humans compare anonymous model answers pairwise and ratings emerge Elo-style, as in chess. Captures conversational preference; systematically overweights style and confidence. Directional signal, not a procurement document.