AAAI 2025technical2 citations

SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses

Dongwei Jiang, Jingyu Zhang, Orion Weller, Nathaniel Weir, Benjamin Van Durme, Daniel Khashabi

Abstract

Can LLMs consistently improve their previous outputs for better results? For this to be true, LLMs would need to be better at discriminating among previously-generated alternatives, than generating initial responses. We explore the validity of this hypothesis in practice. We first formulate a unified framework that allows us to compare the generative and discriminative capability of any model on any task. In our resulting experimental analysis of several open-source and industrial LLMs, we observe that model’s are not reliably better at discriminating among previously-generated alternatives than generating initial responses. This finding challenges the notion that LLMs may be able to enhance their performance only through their own judgment.

BibTeX
@article{Jiang_Zhang_Weller_Weir_Van Durme_Khashabi_2025, title={SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses}, volume={39}, url={https://ojs.aaai.org/index.php/AAAI/article/view/34603}, DOI={10.1609/aaai.v39i23.34603}, abstractNote={Can LLMs consistently improve their previous outputs for better results? For this to be true, LLMs would need to be better at discriminating among previously-generated alternatives, than generating initial responses. We explore the validity of this hypothesis in practice. We first formulate a unified framework that allows us to compare the generative and discriminative capability of any model on any task. In our resulting experimental analysis of several open-source and industrial LLMs, we observe that model’s are not reliably better at discriminating among previously-generated alternatives than generating initial responses. This finding challenges the notion that LLMs may be able to enhance their performance only through their own judgment.}, number={23}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Jiang, Dongwei and Zhang, Jingyu and Weller, Orion and Weir, Nathaniel and Van Durme, Benjamin and Khashabi, Daniel}, year={2025}, month={Apr.}, pages={24266-24275} }
SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses · AAAI 2025