When agreement becomes a failure mode
Why a helpful-sounding answer can still be the wrong answer.
Mrinank Sharma et al.
Towards Understanding Sycophancy in Language Models
The pingpong library
What makes an AI answer worth relying on?
We examine whether AI critique improves an answer and what the findings mean for pingpong.
Our thesis
Each model reviews the work so far against your original question. Its instructions make clear that it does not have to agree with the answer it receives.
Read our approach26 notes
Why a helpful-sounding answer can still be the wrong answer.
Mrinank Sharma et al.
Towards Understanding Sycophancy in Language Models
Resisting pressure and accepting evidence are different behaviors.
Huanhuan Ma et al.
Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update
Capability and resistance to user pressure need separate tests.
Jerry Wei et al.
Simple synthetic data reduces sycophancy in large language models
General benchmark strength can miss a specific conversational weakness.
Ethan Perez et al.
Discovering Language Model Behaviors with Model-Written Evaluations
Early evidence for review between language-model instances.
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch
Improving Factuality and Reasoning in Language Models through Multiagent Debate
The research case for using model outputs as context, not just displaying them side by side.
Junlin Wang, Jue Wang, Ben Athiwaratkun, Ce Zhang, and James Zou
Mixture-of-Agents Enhances Large Language Model Capabilities
Repeated reflection can preserve the mistake it is meant to correct.
Tian Liang et al.
Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
A serious review system must be compared with simpler alternatives.
Andries Smit, Paul Duckworth, Nathan Grinsztajn, Thomas D. Barrett, and Arnu Pretorius
Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs
Repeating a strong model can beat mixing weaker ones.
Wenzhe Li, Yong Lin, Mengzhou Xia, and Chi Jin
Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?
Separate the benefit of multiple samples from the benefit of discussion.
Hyeong Kyu Choi, Xiaojin Zhu, and Yixuan Li
Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?
Feedback can change an output without changing model weights.
Aman Madaan et al.
Self-Refine: Iterative Refinement with Self-Feedback
Revision needs a reason beyond the existence of another pass.
Jie Huang et al.
Large Language Models Cannot Self-Correct Reasoning Yet
A review is valuable when it changes subsequent work.
Noah Shinn et al.
Reflexion: Language Agents with Verbal Reinforcement Learning
A concrete objection can help a human evaluate a polished answer.
William Saunders et al.
Self-critiquing models for assisting human evaluators
External feedback gives criticism something firmer than another opinion.
Zhibin Gou et al.
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
Verification benefits from controlling what the checker sees.
Shehzaad Dhuliawala et al.
Chain-of-Verification Reduces Hallucination in Large Language Models
Disagreement is a diagnostic signal, not a verdict.
Potsawee Manakul, Adian Liusie, and Mark J. F. Gales
SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
Sample several reasoning paths before adding a review protocol.
Xuezhi Wang et al.
Self-Consistency Improves Chain of Thought Reasoning in Language Models
The route to an answer deserves evaluation too.
Hunter Lightman et al.
Let's Verify Step by Step
Length, answer order, and model identity can influence an AI judge.
Lianmin Zheng et al.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Plausible reasoning can rationalize a biased answer.
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman
Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
Long review histories can hide the most important qualification.
Nelson F. Liu et al.
Lost in the Middle: How Language Models Use Long Contexts
Self-assessment depends on how and where it is measured.
Saurav Kadavath et al.
Language Models (Mostly) Know What They Know
A human study offers a useful warning about shared answers.
Jan Lorenz, Heiko Rauhut, Frank Schweitzer, and Dirk Helbing
How social influence can undermine the wisdom of crowd effect
An early proposal connects the interaction protocol to the quality of oversight.
Geoffrey Irving, Paul Christiano, and Dario Amodei
AI safety via debate
The updated Lenz snapshot measures disagreement, not which model is right.
Kosta Jordanov, David Yordanov, and Yana Jordanova
Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks
No notes match this search.