The finding
Liang and colleagues describe a failure in which self-reflection stops producing useful alternatives after a model settles on an answer. Their debate framework improves results on two tested datasets. They also find that the degree of opposition and the stopping policy matter. [1]
What it means for Pingpong
Pingpong's later reviewers are allowed to change the approach, not merely edit wording. That matters when an initial answer solves the wrong problem or accepts a weak premise. A review pass should be able to redirect the work while still answering the user's original question.
The limit
The paper uses a debate and judge protocol on specific tasks. It neither proves that disagreement is always helpful nor that a sequential chain escapes anchoring. Pingpong does not require a reviewer to invent an objection.
Source
Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
Tian Liang et al. (2023)