The finding
Smit and colleagues compare debate strategies with alternatives such as self-consistency. Debate does not reliably win across their settings. Tuning protocol choices, including how readily agents agree, can materially change the results. [1]
What it means for Pingpong
This puts the emphasis in the right place for Pingpong: the review conditions, not the spectacle of several models. The framing at a handoff is a testable part of the system. It should be compared with the same chain without that framing and with a strong single-model baseline.
The limit
A successful example is not enough to justify extra latency or API cost. Matched-budget comparisons are needed before claiming that Pingpong is more efficient or more accurate.
Source
Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs
Andries Smit, Paul Duckworth, Nathan Grinsztajn, Thomas D. Barrett, and Arnu Pretorius (2023)