Applied AI·Evaluation and reliability
you read that a model scores ninety on a public suite, try it on your own task, and find the number told you almost nothing.
Benchmark
Draft summary, pending review
A public, standardised eval suite: MMLU (broad multiple-choice knowledge), HumanEval (short Python problems), SWE-bench (resolving real GitHub issues), MTEB (embeddings), among many. Useful for rough model comparison; none of them measures your task, which is why you build your own set.