The Practitioner's LLM Curriculum ← Week 8 · Judge Agreement Visualizer
Interactive · Week 8 · Section 6

How noisy is your LLM judge?

Three raters grade the same 50 responses: two humans and an LLM-as-judge. Move the subjectivity slider — from "did the response include the correct refund amount?" (factual, binary) to "was this response delightful?" (subjective). Watch how all three raters fall apart together as subjectivity rises, and how the LLM judge falls apart faster than humans. The visceral lesson: LLM-as-judge isn't free, and the noise depends entirely on what you're evaluating.

Task subjectivity 0.50
0 · binary fact-check 1 · subjective quality
Example task at this subjectivity
—
Human A vs truth
—
accuracy
Human B vs truth
—
accuracy
LLM judge vs truth
—
accuracy
Judge below human
—
accuracy gap

Per-sample verdicts · 50 samples per rater

green = matches truth · red = disagrees
Human A moderate-noise rater —
Human B correlated with Human A —
LLM judge independent · noisier —

Pairwise agreement

how often any two raters give the same verdict
Human A vs Human B
—
inter-human agreement (the ceiling)
Human A vs LLM judge
—
judge approximating human A
Human B vs LLM judge
—
judge approximating human B

What's happening here