Research library

Multi-model review Research note

The rules of a review can matter as much as the reviewers

An early proposal connects the interaction protocol to the quality of oversight.

The finding

The authors propose training agents in a debate game judged by a human. Their argument is that competing agents could expose problems a judge would struggle to discover alone. The paper includes a small initial experiment and discusses substantial open questions. [1]

What it means for Pingpong

Pingpong also treats the conditions of review as part of the product. A capable model can be asked to endorse a draft, attack it, or assess it without an obligation to do either. Those are different tasks. Pingpong chooses the last: independent judgment directed toward a useful answer.

The limit

The proposal concerns training incentives and adversarial debate, not Pingpong's prompted sequential review. It is conceptual background, not a safety guarantee or a proof that our system has solved oversight.

Source

AI safety via debate

Geoffrey Irving, Paul Christiano, and Dario Amodei (2018)

Research paper, arXiv archive
DOI: 10.48550/arXiv.1805.00899

Paper PDF

This is Pingpong's interpretation of external research, not a result from a Pingpong experiment or an endorsement by the authors. Editorial standard