Scope
This is a curated starting collection, not a systematic review or an exhaustive bibliography. We selected work directly relevant to critique, revision, sycophancy, shared context, verification, and evaluation. It includes positive results, counterevidence, conceptual work, and one clearly labeled industry report.
Sources were checked on September 8, 2026. Paper years refer to their first public release. A conference label is included where verified; an arXiv link identifies an archive copy, not proof of peer review. Preprints can change. Findings concern the models, tasks, and protocols studied, not every future model.
Findings and interpretation
Each note separates the source's finding, our interpretation for Pingpong, and the limit of that connection. The cited authors have not endorsed Pingpong. Their benchmark results are not our product results. We do not turn model preference scores into claims about real-world decision accuracy.
Notes link to the original research record and DOI where available. We summarize instead of reproducing the paper. The library is published by Pingpong, a product of Adore LLC, and is not an independent review of the product.
The evaluation we would want to see
The questions below describe a proposed evaluation, not experiments completed for this release.
- Does the framing help? Compare the selected chain with and without the review frame, holding model versions, order, inputs, and sampling settings fixed.
- Does the chain justify its budget? Compare with the best single model, repeated-best-model sampling, and simple aggregation at matched cost and latency budgets.
- Are revisions actually corrections? Count incorrect-to-correct changes, correct-to-incorrect changes, unsupported claims, and lost qualifications.
- Can reviewers resist pressure and accept evidence? Pair baseless user pushback with genuine corrections. Evaluate both, rather than rewarding disagreement alone.
- Do important objections survive? Test model order, long inputs, conflicting sources, and facts visible only in an attachment. Record which models had direct source access.
- Can someone else reproduce the result? Publish the task set, rubric, model versions, budgets, failure cases, and uncertainty estimates. Use blinded human assessment where an answer key is unavailable.
Corrections and contributions
We welcome relevant papers, corrections to these notes, and proposals for a reproducible evaluation. Include the source link and the specific claim or comparison you want us to examine.