The finding
CRITIC lets a model use tools to check aspects of an initial output before revising it. The authors report gains across question answering, mathematical program synthesis, and toxicity reduction, and emphasize the importance of external feedback. [1]
What it means for Pingpong
Pingpong combines review with source access where a selected provider supports it. That makes evidence quality relevant throughout the chain: a later model should examine the basis of a claim, not inherit certainty from an earlier answer. A retrieved source can still be weak, outdated, or misread.
The limit
Pingpong does not run the CRITIC protocol or verify every claim with a tool. A model's critique is not equivalent to a successful test, a primary source, or a confirmed observation.
Source
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
Zhibin Gou et al. (2023)