Applied AI·Evaluation and reliability
you have a second model score the first one's output against a rubric, because there is no exact answer to diff against.
LLM-as-judge
Draft summary, pending review
Using a strong model to score outputs against a rubric when correctness is fuzzy: answer quality, faithfulness to sources, tone. Powerful and biased; calibrate the judge against a sample of human labels before trusting its numbers, and keep the judge model pinned.