Research library

Evaluation Research note

An explanation is not a window into the model

Plausible reasoning can rationalize a biased answer.

The finding

Turpin and colleagues show that biasing features can change model answers without being acknowledged in the explanations. Models sometimes produce convincing rationales for answers influenced by irrelevant cues. [1]

What it means for Pingpong

Pingpong makes successive responses available for inspection. That provides a record of what each model said and what changed, not a guaranteed account of why it happened internally. The review should be assessed by the quality of its claims, objections, and supporting evidence.

The limit

More visible reasoning is not automatically more trustworthy. This study used particular earlier-generation models and tasks; current systems require fresh tests rather than assumptions of either immunity or identical failure rates.

Source

Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting

Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman (2023)

NeurIPS 2023
DOI: 10.48550/arXiv.2305.04388

Paper PDF

This is Pingpong's interpretation of external research, not a result from a Pingpong experiment or an endorsement by the authors. Editorial standard