The Practitioner's LLM Curriculum ← Week 8 · Judge Agreement Visualizer
Interactive · Week 8 · Section 6

How noisy is your LLM judge?

Three raters grade the same 50 responses: two humans and an LLM-as-judge. Move the subjectivity slider — from "did the response include the correct refund amount?" (factual, binary) to "was this response delightful?" (subjective). Watch how all three raters fall apart together as subjectivity rises, and how the LLM judge falls apart faster than humans. The visceral lesson: LLM-as-judge isn't free, and the noise depends entirely on what you're evaluating.

Task subjectivity 0.50
0 · binary fact-check 1 · subjective quality
Example task at this subjectivity
Human A vs truth
accuracy
Human B vs truth
accuracy
LLM judge vs truth
accuracy
Judge below human
accuracy gap

Per-sample verdicts · 50 samples per rater

green = matches truth · red = disagrees
Human A moderate-noise rater
Human B correlated with Human A
LLM judge independent · noisier

Pairwise agreement

how often any two raters give the same verdict
Human A vs Human B
inter-human agreement (the ceiling)
Human A vs LLM judge
judge approximating human A
Human B vs LLM judge
judge approximating human B

What's happening here