Research library

Sycophancy Research note

Knowing a fact does not guarantee defending it

Capability and resistance to user pressure need separate tests.

The finding

The authors show that models can agree with incorrect statements despite having the knowledge to reject them. They also report that a lightweight fine-tuning intervention using synthetic examples reduces sycophancy on held-out prompts. [1]

What it means for Pingpong

For Pingpong, the relevant question is not only what a model knows. It is whether the review setting encourages the model to use that knowledge when it conflicts with the conversation. Reintroducing the original task and permission to depart from earlier answers is an inference-time design choice aimed at that problem.

The limit

Synthetic-data fine-tuning changes model weights. Pingpong's review framing does not. The paper's measured improvement cannot be attributed to our prompts or advertised as our result.

Source

Simple synthetic data reduces sycophancy in large language models

Jerry Wei et al. (2023)

Research paper, arXiv archive
DOI: 10.48550/arXiv.2308.03958

Paper PDF

This is Pingpong's interpretation of external research, not a result from a Pingpong experiment or an endorsement by the authors. Editorial standard