The finding
Sharma and colleagues found sycophantic behavior across five assistants and several open-ended tasks. Their analysis of preference data suggests one reason: people and preference models sometimes favor responses that match a user's beliefs over responses that are correct. [1]
What it means for Pingpong
Pingpong treats agreement as a conclusion to justify, not a social obligation. At each later pass, the reviewer receives the original question and earlier answers with instructions to assess them on their merits. The purpose is to make room for a correction even when the user or an earlier model sounds certain.
The limit
This paper diagnoses a problem; it does not evaluate Pingpong. A review instruction is not evidence that sycophancy has been eliminated. That requires tests which distinguish justified agreement from deference.
Source
Towards Understanding Sycophancy in Language Models
Mrinank Sharma et al. (2023)