The finding
Li and colleagues find that aggregating several outputs from one strong model can outperform mixtures of different models. Their analysis identifies a trade-off between diversity and the quality of the candidate answers. Greater diversity can help when quality is controlled. [1]
What it means for Pingpong
Different provider names are not, by themselves, a reason to trust Pingpong. The rationale for a mixed chain is the possibility of useful differences in assessment while retaining strong reviewers. Whether that benefit survives a particular task, ordering, and budget is an empirical question.
The limit
This is counterevidence to a blanket claim that more model families always mean better answers. A repeated-best-model baseline belongs in any credible evaluation of Pingpong.
Source
Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?
Wenzhe Li, Yong Lin, Mengzhou Xia, and Chi Jin (2025)