Research library

Sycophancy Research note

Test the behavior the product depends on

General benchmark strength can miss a specific conversational weakness.

The finding

Perez and colleagues use language models to generate evaluation datasets, with human checks of relevance and labels. The evaluations reveal behaviors, including sycophancy, that can worsen with model scale or additional human-feedback training in the tested settings. [1]

What it means for Pingpong

A critique product needs critique-specific tests. For Pingpong, that means questions with false premises, confident but unsupported prior answers, and cases where agreement is warranted. Testing only whether the final prose is fluent would miss the behavior the review chain is supposed to improve.

The limit

Generated evaluations can contain errors or reflect the generator's own blind spots. Human review, held-out examples, and documented rubrics remain necessary. This is a basis for an evaluation program, not a completed Pingpong benchmark.

Source

Discovering Language Model Behaviors with Model-Written Evaluations

Ethan Perez et al. (2022)

Research paper, arXiv archive
DOI: 10.48550/arXiv.2212.09251

Paper PDF

This is Pingpong's interpretation of external research, not a result from a Pingpong experiment or an endorsement by the authors. Editorial standard