jargon

Applied AI·Evaluation and reliability

a model posts a startling score on a public test set and does no better than the last one on yours, because the answers were in its training data.

Benchmark contamination

Draft summary, pending review

Benchmark answers leaking into training data, inflating scores without real capability. A known chronic problem; treat dramatic public numbers with suspicion and weight private, held-out evals accordingly.