The finding
Du and colleagues ask model instances to propose answers and discuss them over multiple rounds. They report improvements on selected mathematical, strategic-reasoning, and factual tasks. The work demonstrates that interaction between generated answers can be useful, rather than merely adding more text. [1]
What it means for Pingpong
Pingpong uses a sequential review chain, not a round-table debate. Each later model can inspect previous responses and produce a revised answer to the original question. The shared idea is that another assessment can expose a weakness the initial response left intact.
The limit
The study's models, tasks, and debate protocol differ from Pingpong. Its gains do not establish that five different providers, a particular order, or our framing outperform other methods.
Source
Improving Factuality and Reasoning in Language Models through Multiagent Debate
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch (2023)