The finding
Across seven benchmarks, the authors find that majority voting explains much of the gain often attributed to debate. Their theoretical model also shows why peer updates need not improve expected correctness under its assumptions. Targeted interventions can change those dynamics. [1]
What it means for Pingpong
Pingpong's thesis concerns the quality of revision, not how quickly models converge. A later answer should preserve valid objections and correct identifiable errors. Agreement alone is a poor success metric for a system intended to question a premise or an earlier response.
The limit
The theorem is conditional on a formal model, not a universal impossibility result. Equally, it cannot be used to claim that Pingpong's review framing is a proven corrective intervention. That comparison remains to be tested.
Source
Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?
Hyeong Kyu Choi, Xiaojin Zhu, and Yixuan Li (2025)