The finding
The authors show that models can agree with incorrect statements despite having the knowledge to reject them. They also report that a lightweight fine-tuning intervention using synthetic examples reduces sycophancy on held-out prompts. [1]
What it means for Pingpong
For Pingpong, the relevant question is not only what a model knows. It is whether the review setting encourages the model to use that knowledge when it conflicts with the conversation. Reintroducing the original task and permission to depart from earlier answers is an inference-time design choice aimed at that problem.
The limit
Synthetic-data fine-tuning changes model weights. Pingpong's review framing does not. The paper's measured improvement cannot be attributed to our prompts or advertised as our result.
Source
Simple synthetic data reduces sycophancy in large language models
Jerry Wei et al. (2023)