DPO's loss is a one-line function of the margin between chosen and rejected log-probabilities. The β parameter controls how steeply the loss falls as the margin grows — and how aggressively the gradient saturates. Drag β to see the curve change shape, run training steps, and watch how four preference pairs converge at different rates because of where they started on the curve.