Research library

Evaluation Research note

The evaluator can prefer the wrong things

Length, answer order, and model identity can influence an AI judge.

The finding

The paper shows that strong model judges can agree substantially with human preferences, while documenting position, verbosity, and self-enhancement biases. It also distinguishes these preference evaluations from traditional capability benchmarks. [1]

What it means for Pingpong

An AI judge can be useful in reviewing Pingpong outputs, but it cannot be the only authority on quality. Evaluations should hide model identity, vary answer order, and use a task-specific rubric. A longer or more polished final response should not win simply because it looks more complete.

The limit

Agreement with human preference is not the same as truth. Factual claims need evidence, and consequential recommendations need evaluation by people with relevant expertise.

Source

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Lianmin Zheng et al. (2023)

NeurIPS 2023 Datasets and Benchmarks
DOI: 10.48550/arXiv.2306.05685

Paper PDF

This is Pingpong's interpretation of external research, not a result from a Pingpong experiment or an endorsement by the authors. Editorial standard