LLM-as-judge uses a strong language model to score another model's outputs against a rubric, rating accuracy, helpfulness, groundedness, or custom criteria, ena
Using a strong language model to score another model's outputs against a rubric, rating accuracy, helpfulness, groundedness, or custom criteria, which enables evaluation at scales human review cannot reach.
Done casually, judges silently reward their own quirks, like verbosity preference or self-preference, which is how teams ship regressions with green dashboards. Professional use calibrates judges against human ratings, audits for bias, and gives precise rubrics.