Research library

Evaluation Research note

A correct conclusion can conceal a broken step

The route to an answer deserves evaluation too.

The finding

Lightman and colleagues compare training with feedback on intermediate reasoning steps against feedback on final outcomes. In their MATH experiments, process supervision performs better. They release a large dataset of human-labeled reasoning steps. [1]

What it means for Pingpong

The lesson for Pingpong is to evaluate substantive changes between passes, not just whether the final paragraph sounds good. Did a reviewer repair an invalid inference? Did it remove a needed qualification? Retaining the sequence makes those questions possible to investigate.

The limit

Process supervision is a training method. Pingpong's inference-time review is not a process-supervised reward model. A visible response history also does not provide direct access to a model's internal computation.

Source

Let's Verify Step by Step

Hunter Lightman et al. (2023)

Research paper, arXiv archive
DOI: 10.48550/arXiv.2305.20050

Paper PDF

This is Pingpong's interpretation of external research, not a result from a Pingpong experiment or an endorsement by the authors. Editorial standard