The finding
Huang and colleagues study intrinsic self-correction: revision without external feedback. On the reasoning tasks they evaluate, models struggle to correct themselves and sometimes become less accurate after another attempt. [1]
What it means for Pingpong
Pingpong brings another model into the review, but that model is not an answer key. The reason to retain the original question and earlier responses is to let a reviewer identify a specific gap, not to assume the latest wording must be superior. A good chain must be allowed to preserve a correct answer.
The limit
The title describes the models and conditions studied, not an eternal limit on AI. For Pingpong, the important risk is regression: tests must count correct-to-incorrect revisions as well as successful corrections.
Source
Large Language Models Cannot Self-Correct Reasoning Yet
Jie Huang et al. (2023)