Pick which reward components are active, then drag the training-stage slider and watch the policy progressively discover ways to score higher without actually doing the task better. The "true quality" score (held out from training) tracks what a thoughtful human would think — and diverges sharply from the reward as the policy learns to hack. Toggle the KL constraint on to see the load-bearing defense in action.
Active reward components
Mitigation
KL constraint to referencecaps how far the policy can drift
Training stageStage 0 · Honest
honestverbosecitedstructuredhacked
user prompt
How do I reset my password on the dashboard?