The finding
Lightman and colleagues compare training with feedback on intermediate reasoning steps against feedback on final outcomes. In their MATH experiments, process supervision performs better. They release a large dataset of human-labeled reasoning steps. [1]
What it means for Pingpong
The lesson for Pingpong is to evaluate substantive changes between passes, not just whether the final paragraph sounds good. Did a reviewer repair an invalid inference? Did it remove a needed qualification? Retaining the sequence makes those questions possible to investigate.
The limit
Process supervision is a training method. Pingpong's inference-time review is not a process-supervised reward model. A visible response history also does not provide direct access to a model's internal computation.
Source
Hunter Lightman et al. (2023)