2024
Every Answer Matters: Evaluating Commonsense with Probabilistic Measures
ACL 2024long
Large language models have demonstrated impressive performance on commonsense tasks; however, these tasks are often posed as multiple-choice questions, allowing models to exploit systematic biases. Commonsense is also inherently probabilistic with multiple correct answers. The purpose of “boiling wa…