The finding
The authors propose training agents in a debate game judged by a human. Their argument is that competing agents could expose problems a judge would struggle to discover alone. The paper includes a small initial experiment and discusses substantial open questions. [1]
What it means for Pingpong
Pingpong also treats the conditions of review as part of the product. A capable model can be asked to endorse a draft, attack it, or assess it without an obligation to do either. Those are different tasks. Pingpong chooses the last: independent judgment directed toward a useful answer.
The limit
The proposal concerns training incentives and adversarial debate, not Pingpong's prompted sequential review. It is conceptual background, not a safety guarantee or a proof that our system has solved oversight.
Source
Geoffrey Irving, Paul Christiano, and Dario Amodei (2018)