The finding
Kadavath and colleagues find encouraging self-evaluation and calibration results when questions are presented in suitable formats. They also report difficulty generalizing calibration to new tasks. [1]
What it means for Pingpong
Pingpong should leave room for a reviewer to say that an answer lacks support. The useful outcome is not a mandatory conclusion at any cost. It may be a narrower recommendation, a missing assumption, or a clear statement that more evidence is needed.
The limit
A model saying it is confident or uncertain is not a calibrated probability for a real decision. Confidence needs validation on the task at hand. This paper does not justify displaying an untested reliability score for Pingpong.
Source
Language Models (Mostly) Know What They Know
Saurav Kadavath et al. (2022)