The finding
Lenz asks five frontier models to assess 1,000 user-submitted claims. In the revised v1.1 report, 997 receive usable verdicts from every model; 63% lack unanimity. The report examines differences in verdicts and confidence, not accuracy against a human-labeled answer key. [1]
What it means for pingpong
A model split is a reason to investigate the question, evidence, or definitions. pingpong takes a different approach from this study's separate assessments: it lets later models examine prior responses and revise the answer. The goal is to turn an unresolved objection into useful work, not hide it behind a vote.
The limit
This is an industry research report, not a peer-reviewed validation of pingpong. The September 8 note replaces our earlier discussion of v1.0; the old figures and informal examples are not evidence for the revised dataset.
Source
Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks
Kosta Jordanov, David Yordanov, and Yana Jordanova (2026)