pingpong Research Sources checked 2026-09-08. Commentary and scope: https://pingpongit.com/research/method/ Mrinank Sharma et al. (2023). Towards Understanding Sycophancy in Language Models. ICLR 2024. https://arxiv.org/abs/2310.13548. DOI: 10.48550/arXiv.2310.13548. Huanhuan Ma et al. (2026). Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update. Preprint; EMNLP 2026 Findings acceptance reported by authors. https://arxiv.org/abs/2608.26511. DOI: 10.48550/arXiv.2608.26511. Jerry Wei et al. (2023). Simple synthetic data reduces sycophancy in large language models. https://arxiv.org/abs/2308.03958. DOI: 10.48550/arXiv.2308.03958. Ethan Perez et al. (2022). Discovering Language Model Behaviors with Model-Written Evaluations. https://arxiv.org/abs/2212.09251. DOI: 10.48550/arXiv.2212.09251. Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch (2023). Improving Factuality and Reasoning in Language Models through Multiagent Debate. ICML 2024. https://arxiv.org/abs/2305.14325. DOI: 10.48550/arXiv.2305.14325. Junlin Wang, Jue Wang, Ben Athiwaratkun, Ce Zhang, and James Zou (2024). Mixture-of-Agents Enhances Large Language Model Capabilities. https://arxiv.org/abs/2406.04692. DOI: 10.48550/arXiv.2406.04692. Tian Liang et al. (2023). Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate. EMNLP 2024. https://arxiv.org/abs/2305.19118. DOI: 10.48550/arXiv.2305.19118. Andries Smit, Paul Duckworth, Nathan Grinsztajn, Thomas D. Barrett, and Arnu Pretorius (2023). Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs. https://arxiv.org/abs/2311.17371. DOI: 10.48550/arXiv.2311.17371. Wenzhe Li, Yong Lin, Mengzhou Xia, and Chi Jin (2025). Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?. https://arxiv.org/abs/2502.00674. DOI: 10.48550/arXiv.2502.00674. Hyeong Kyu Choi, Xiaojin Zhu, and Yixuan Li (2025). Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?. https://arxiv.org/abs/2508.17536. DOI: 10.48550/arXiv.2508.17536. Aman Madaan et al. (2023). Self-Refine: Iterative Refinement with Self-Feedback. https://arxiv.org/abs/2303.17651. DOI: 10.48550/arXiv.2303.17651. Jie Huang et al. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. ICLR 2024. https://arxiv.org/abs/2310.01798. DOI: 10.48550/arXiv.2310.01798. Noah Shinn et al. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. https://arxiv.org/abs/2303.11366. DOI: 10.48550/arXiv.2303.11366. William Saunders et al. (2022). Self-critiquing models for assisting human evaluators. https://arxiv.org/abs/2206.05802. DOI: 10.48550/arXiv.2206.05802. Zhibin Gou et al. (2023). CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing. https://arxiv.org/abs/2305.11738. DOI: 10.48550/arXiv.2305.11738. Shehzaad Dhuliawala et al. (2023). Chain-of-Verification Reduces Hallucination in Large Language Models. https://arxiv.org/abs/2309.11495. DOI: 10.48550/arXiv.2309.11495. Potsawee Manakul, Adian Liusie, and Mark J. F. Gales (2023). SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. EMNLP 2023. https://arxiv.org/abs/2303.08896. DOI: 10.48550/arXiv.2303.08896. Xuezhi Wang et al. (2022). Self-Consistency Improves Chain of Thought Reasoning in Language Models. ICLR 2023. https://arxiv.org/abs/2203.11171. DOI: 10.48550/arXiv.2203.11171. Hunter Lightman et al. (2023). Let's Verify Step by Step. https://arxiv.org/abs/2305.20050. DOI: 10.48550/arXiv.2305.20050. Lianmin Zheng et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023 Datasets and Benchmarks. https://arxiv.org/abs/2306.05685. DOI: 10.48550/arXiv.2306.05685. Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. NeurIPS 2023. https://arxiv.org/abs/2305.04388. DOI: 10.48550/arXiv.2305.04388. Nelson F. Liu et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. TACL acceptance reported in arXiv record. https://arxiv.org/abs/2307.03172. DOI: 10.48550/arXiv.2307.03172. Saurav Kadavath et al. (2022). Language Models (Mostly) Know What They Know. https://arxiv.org/abs/2207.05221. DOI: 10.48550/arXiv.2207.05221. Jan Lorenz, Heiko Rauhut, Frank Schweitzer, and Dirk Helbing (2011). How social influence can undermine the wisdom of crowd effect. PNAS 108(22), 9020-9025. https://pmc.ncbi.nlm.nih.gov/articles/PMC3107299/. DOI: 10.1073/pnas.1008636108. Geoffrey Irving, Paul Christiano, and Dario Amodei (2018). AI safety via debate. https://arxiv.org/abs/1805.00899. DOI: 10.48550/arXiv.1805.00899. Kosta Jordanov, David Yordanov, and Yana Jordanova (2026). Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks. Lenz Research report, v1.1, August 7, 2026. https://lenz.io/research/llm-disagreement. DOI: 10.5281/zenodo.21829261.