The finding
Self-Refine uses the same model to generate an answer, give feedback, and revise it. Across seven tasks, the authors report improvements over one-step generation using human assessments and automatic metrics, without additional training. [1]
What it means for Pingpong
The relevance to Pingpong is the value of allocating a separate opportunity to review. The product uses later models for that opportunity and supplies a fresh frame for evaluating earlier work. Its final answer is the result of the selected sequence, rather than a page of unrelated first drafts.
The limit
This study supports iterative revision in its tested settings. It does not establish that replacing the self-reviewer with another model is always better, or that each additional pass improves correctness.
Source
Self-Refine: Iterative Refinement with Self-Feedback
Aman Madaan et al. (2023)