The finding
Perez and colleagues use language models to generate evaluation datasets, with human checks of relevance and labels. The evaluations reveal behaviors, including sycophancy, that can worsen with model scale or additional human-feedback training in the tested settings. [1]
What it means for Pingpong
A critique product needs critique-specific tests. For Pingpong, that means questions with false premises, confident but unsupported prior answers, and cases where agreement is warranted. Testing only whether the final prose is fluent would miss the behavior the review chain is supposed to improve.
The limit
Generated evaluations can contain errors or reflect the generator's own blind spots. Human review, held-out examples, and documented rubrics remain necessary. This is a basis for an evaluation program, not a completed Pingpong benchmark.
Source
Discovering Language Model Behaviors with Model-Written Evaluations
Ethan Perez et al. (2022)