A synthetic 100-query eval set. Each query has a hidden retrieval difficulty and a hidden generation difficulty. Move the sliders to set how capable your retrieval and generation are. Watch how end-to-end correctness emerges from the product of the two — and how the bottleneck shifts as you improve one without the other. Most teams overinvest in the easier-to-improve metric while the bottleneck silently caps performance. The visceral lesson: improving the wrong metric doesn't move the user-visible number.