Three raters grade the same 50 responses: two humans and an LLM-as-judge. Move the subjectivity slider — from "did the response include the correct refund amount?" (factual, binary) to "was this response delightful?" (subjective). Watch how all three raters fall apart together as subjectivity rises, and how the LLM judge falls apart faster than humans. The visceral lesson: LLM-as-judge isn't free, and the noise depends entirely on what you're evaluating.