Research library

Verification Research note

Frontier models can disagree on the same factual claim

The updated Lenz snapshot measures disagreement, not which model is right.

The finding

Lenz asks five frontier models to assess 1,000 user-submitted claims. In the revised v1.1 report, 997 receive usable verdicts from every model; 63% lack unanimity. The report examines differences in verdicts and confidence, not accuracy against a human-labeled answer key. [1]

What it means for pingpong

A model split is a reason to investigate the question, evidence, or definitions. pingpong takes a different approach from this study's separate assessments: it lets later models examine prior responses and revise the answer. The goal is to turn an unresolved objection into useful work, not hide it behind a vote.

The limit

This is an industry research report, not a peer-reviewed validation of pingpong. The September 8 note replaces our earlier discussion of v1.0; the old figures and informal examples are not evidence for the revised dataset.

Source

Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks

Kosta Jordanov, David Yordanov, and Yana Jordanova (2026)

Lenz Research report, v1.1, August 7, 2026
DOI: 10.5281/zenodo.21829261

This is pingpong's interpretation of external research, not a result from a pingpong experiment or an endorsement by the authors. Editorial standard