Articles

What Is LLM-as-a-Judge?

LLM-as-a-judge uses one language model to score another model's outputs against a rubric, replacing slow human review for large-scale evaluation.

Chisato Chisato · · 4 min read
An abstract purple render of neural network fibers

LLM-as-a-judge is an evaluation technique that uses one large language model to score, rank, or critique the outputs of another model — or the same model’s own outputs — instead of relying on a human reviewer for every judgment. A judge model is given the original prompt, the output being evaluated, and a rubric or reference answer, and returns a score or a verdict, at a speed and cost no human review process can match.

Why evaluation needs this at all

Judging whether a model’s output is good is often not something a simple string match or unit test can settle. Ask a model to summarize an article, write a product description, or answer an open-ended support question, and there’s no single correct string to compare against — quality is a judgment call about accuracy, tone, completeness, and relevance. This is the same problem LLM evals run into broadly: for open-ended generation, the tests that scale (exact match, regex, unit tests) can’t capture what “good” means, and the tests that can capture it — a human reading every output — don’t scale.

LLM-as-a-judge sits between those two extremes. It can’t replace a careful human review of edge cases, but it can apply a reasonably consistent rubric across thousands of outputs, cheaply enough to run on every pull request, every model version, or every day’s production traffic — the kind of volume no human review team can sustain.

How a judge prompt is built

A working LLM-as-a-judge setup usually needs more than “is this good?” A typical judge prompt includes:

  1. The original input the model being evaluated was responding to.
  2. The output being judged.
  3. A rubric or reference answer — explicit criteria (“does the summary mention the article’s main conclusion?”, “is the code free of syntax errors?”) or a known-good example to compare against.
  4. A required output format, usually a score on a fixed scale or a structured verdict, so results can be aggregated programmatically. This is where structured outputs or a forced tool call earn their keep — a judge that has to return {"score": 4, "reasoning": "..."} is far easier to build a pipeline around than one that free-writes a paragraph.

Two common judging patterns show up repeatedly. Pointwise scoring judges a single output in isolation against a rubric, producing an absolute score. Pairwise comparison shows the judge two outputs — often from two different model versions or two different prompts — and asks which is better, which tends to be an easier and more reliable judgment for a model to make than assigning an absolute number, since it’s a relative call rather than a calibration exercise.

Where it breaks down

LLM-as-a-judge inherits the same failure modes as the model doing the judging, and adds a few of its own:

  • Self-preference bias. Judge models measurably tend to rate outputs from their own model family more favorably, even when quality is genuinely comparable — a confound that matters a lot if you’re using a judge to compare your model against a competitor’s.
  • Position bias in pairwise comparisons. Some judges are more likely to prefer whichever answer is shown first (or second), regardless of content, so serious pairwise setups run each comparison both ways and check for order flips.
  • Length bias. Judges have a documented tendency to rate longer answers as better, independent of whether the extra length adds anything — a rubric that doesn’t explicitly guard against this can end up optimizing models toward verbosity.
  • The judge can hallucinate its own verdict. A judge that’s asked to check a factual claim can be just as confidently wrong as the model it’s evaluating; see LLM hallucinations explained for why fluent confidence and correctness aren’t the same thing in either direction.

The standard mitigation for most of these is calibration: periodically checking judge verdicts against a smaller set of human-labeled examples, using a stronger or differently-sourced model as the judge than the one being evaluated, and keeping rubrics narrow and specific rather than open-ended (“did the response include X, Y, and Z?” rather than “was this a good response?”).

LLM-as-a-judge vs traditional evals

Rule-based evals (exact match, unit tests)LLM-as-a-judge
Works well forDeterministic tasks with a clear right answerOpen-ended generation, tone, reasoning quality
ConsistencyPerfectly consistentCan drift, and carries model-specific biases
Setup costWrite test cases and assertionsWrite and calibrate a rubric prompt
Scales toAny volume, near-zero marginal costHigh volume, but real per-call cost
Human involvementNone after test authoringSpot-checking to catch judge drift

In practice, most serious evaluation pipelines use both: rule-based checks for anything that has a clean pass/fail definition, and an LLM judge for the open-ended remainder, with periodic human spot-checks keeping the judge itself honest.

The takeaway

LLM-as-a-judge fills the gap between evaluation that scales and evaluation that captures real quality — using a model to apply a rubric across volumes no human team could review, at the cost of inheriting the judging model’s own biases toward self-similarity, position, and length. Treat judge scores as a strong, cheap signal to run continuously, not a ground truth to trust blindly, and calibrate them against human review often enough to catch drift before it skews everything downstream.

Chisato Chisato · · 4 min read

Precision vs Recall, Explained

Precision measures how many of a model's positive predictions were correct; recall measures how many actual positives it found. Why you can't max both.

#AI #Machine Learning #LLMs