jargon

Applied AI·Evaluation and reliability

you have a second model score the first one's output against a rubric, because there is no exact answer to diff against.

LLM-as-judge

Draft summary, pending review

Using a strong model to score outputs against a rubric when correctness is fuzzy: answer quality, faithfulness to sources, tone. Powerful and biased; calibrate the judge against a sample of human labels before trusting its numbers, and keep the judge model pinned.

Commonly confused with