New · France 2027 presidential election: what the candidates propose on AI, quoted and sourced. Explore the tracker →

Last reviewed:

What is LLM evaluation? Definition of evals and LLM as a judge

LLM evaluation (“evals”) means measuring the quality of a model's or AI assistant's answers on a set of representative cases, against success criteria set in advance. LLM as a judge is one of its methods: a second model grades the first one's answers, which makes evaluation possible at scale, at the cost of known biases.

A public leaderboard (benchmark) says little about your use case. To decide whether an AI assistant can go into production, you need an evaluation specific to your business. It rests on three elements. First, precise, measurable success criteria: accuracy, tone, absence of personal data, latency, cost. Anthropic advises against vague goals such as “good performance”. Second, a test set that mirrors the real distribution of requests, edge cases included. Third, a grading method: automatic comparison with the expected answer where possible, human review of a sample, or grading by another model. This last method, LLM as a judge, was studied by Zheng et al. (2023): GPT-4 used as a judge matches human preferences in over 80% of cases, the same level of agreement observed between humans. The authors also describe three biases: position bias (preferring the answer presented first), verbosity bias (preferring the longest) and self-enhancement bias (preferring its own answers). Hence Anthropic's recommendation: use a different model to grade than the one that generates. Evals do not stop at launch. They are rerun at every change of model, prompt or data. In April 2025, OpenAI had to withdraw a GPT-4o update deemed too sycophantic: its offline evals and A/B tests looked good, but no deployment eval measured sycophancy.

Concrete example

Illustrative case: a French regional health insurer is preparing an assistant that answers members' questions about their coverage. Before launch, the team gathers 400 real questions from the customer service history, including 60 trick cases (cancelled coverage, medical questions, out-of-scope requests). Two case handlers write the expected answer for each one. The assistant's answers are graded three ways: automatic checking of reimbursement amounts, grading of clarity and tone by a model from another provider, human review of a random 10% of cases to check that judge. Criterion set by management: zero amount errors and 95% of answers rated correct. The first version reaches 88%. Three iterations on source documents and instructions are needed before launch. The same test set is rerun at every model update.

Comparison

Three grading methods to evaluate an AI assistant
Automatic checkingHuman reviewLLM as a judge
PrincipleCompare with an expected answer (amount, category, format)A business expert grades each answerA second model grades against a rubric
CostVery lowHighLow
SpeedImmediateSlowFast, at scale
Suited toFigures, classifications, extractionsSensitive cases, initial calibrationClarity, tone, relevance, completeness
LimitationUnusable on open-ended answersFew cases coveredPosition, verbosity and self-enhancement biases

FAQ

What is an eval in AI?

An eval is a structured test: a model or AI assistant is given a series of representative cases, and its answers are compared with criteria set in advance (expected answer, business rule, quality rubric). It is the AI equivalent of software acceptance testing.

What is LLM as a judge?

It is a method where a language model grades another model's answers against a rubric (accuracy, clarity, tone). It makes it possible to evaluate thousands of answers quickly. Zheng et al. (2023) showed that a judge such as GPT-4 matches human preferences in over 80% of cases.

What are the biases of LLM as a judge?

Three documented biases: position bias (the judge favours the answer presented first), verbosity bias (it prefers the longest answer) and self-enhancement bias (it favours answers from its own model). Remedies: swap the order, use a judge from another provider, check a sample by hand.

What is the difference between a benchmark and a business evaluation?

A benchmark is a public, standardised test (maths, code, general knowledge) that compares models with each other. A business evaluation measures quality on your own cases, with your criteria. A well-ranked model can fail on your documents, your vocabulary or your rules.

How many test cases are needed to evaluate an AI assistant?

There is no official figure. Anthropic recommends favouring volume with automated grading over a few hand-graded cases. In practice, a few hundred real cases, including a share of trick cases, already give a usable measure for a go-live decision.

Is evaluating an AI assistant mandatory?

Not in general. For high-risk systems, the EU AI Act requires an appropriate level of accuracy and robustness, which must be demonstrated; since the AI Digital Omnibus (Regulation 2026/1744), these obligations apply from December 2, 2027 for Annex III. For other uses, it is a matter of risk control and of evidence in case of an incident.

See also

Further reading

Define success criteria and build evaluations, Claude documentation, Anthropic (external resource)

Sources

  1. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Zheng et al., NeurIPS 2023, arXiv. https://arxiv.org/abs/2306.05685 (accessed 2026-09-30)
  2. Create strong empirical evaluations (success criteria, grading methods), Claude documentation, Anthropic. https://platform.claude.com/docs/en/test-and-evaluate/develop-tests (accessed 2026-09-30)
  3. Expanding on what we missed with sycophancy, OpenAI, May 2, 2025. https://openai.com/index/expanding-on-sycophancy/ (accessed 2026-09-30)
  4. Regulation (EU) 2026/1744 (AI Digital Omnibus), EUR-Lex. https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX%3A32026R1744 (accessed 2026-09-30)

← Back to glossary

Address copied