Research library

Multi-model review Research note

More discussion has to earn its cost

A serious review system must be compared with simpler alternatives.

The finding

Smit and colleagues compare debate strategies with alternatives such as self-consistency. Debate does not reliably win across their settings. Tuning protocol choices, including how readily agents agree, can materially change the results. [1]

What it means for Pingpong

This puts the emphasis in the right place for Pingpong: the review conditions, not the spectacle of several models. The framing at a handoff is a testable part of the system. It should be compared with the same chain without that framing and with a strong single-model baseline.

The limit

A successful example is not enough to justify extra latency or API cost. Matched-budget comparisons are needed before claiming that Pingpong is more efficient or more accurate.

Source

Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs

Andries Smit, Paul Duckworth, Nathan Grinsztajn, Thomas D. Barrett, and Arnu Pretorius (2023)

Research paper, arXiv archive
DOI: 10.48550/arXiv.2311.17371

Paper PDF

This is Pingpong's interpretation of external research, not a result from a Pingpong experiment or an endorsement by the authors. Editorial standard