Research library

Multi-model review Research note

Agreement can improve without the reasoning improving

Separate the benefit of multiple samples from the benefit of discussion.

The finding

Across seven benchmarks, the authors find that majority voting explains much of the gain often attributed to debate. Their theoretical model also shows why peer updates need not improve expected correctness under its assumptions. Targeted interventions can change those dynamics. [1]

What it means for Pingpong

Pingpong's thesis concerns the quality of revision, not how quickly models converge. A later answer should preserve valid objections and correct identifiable errors. Agreement alone is a poor success metric for a system intended to question a premise or an earlier response.

The limit

The theorem is conditional on a formal model, not a universal impossibility result. Equally, it cannot be used to claim that Pingpong's review framing is a proven corrective intervention. That comparison remains to be tested.

Source

Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?

Hyeong Kyu Choi, Xiaojin Zhu, and Yixuan Li (2025)

Research paper, arXiv archive
DOI: 10.48550/arXiv.2508.17536

Paper PDF

This is Pingpong's interpretation of external research, not a result from a Pingpong experiment or an endorsement by the authors. Editorial standard