Research library

Multi-model review Research note

Review needs room to depart from the first answer

Repeated reflection can preserve the mistake it is meant to correct.

The finding

Liang and colleagues describe a failure in which self-reflection stops producing useful alternatives after a model settles on an answer. Their debate framework improves results on two tested datasets. They also find that the degree of opposition and the stopping policy matter. [1]

What it means for Pingpong

Pingpong's later reviewers are allowed to change the approach, not merely edit wording. That matters when an initial answer solves the wrong problem or accepts a weak premise. A review pass should be able to redirect the work while still answering the user's original question.

The limit

The paper uses a debate and judge protocol on specific tasks. It neither proves that disagreement is always helpful nor that a sequential chain escapes anchoring. Pingpong does not require a reviewer to invent an objection.

Source

Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

Tian Liang et al. (2023)

EMNLP 2024
DOI: 10.48550/arXiv.2305.19118

Paper PDF

This is Pingpong's interpretation of external research, not a result from a Pingpong experiment or an endorsement by the authors. Editorial standard