Truthfulness Despite Weak Supervision: Evaluating and Training LLMs Using Peer Prediction
The evaluation and post-training of large language models (LLMs) rely on supervision, but strong supervision for difficult tasks is often unavailable, especially when evaluating strong models. In such cases, models have been demonstrated to exploit evaluation schemes built on such imperfect supervis…