EMNLP 2023short findings0 citations

Decoding Stumpers: Large Language Models vs. Human Problem-Solvers

Alon Goldstein, Miriam Havin, Roi Reichart, Ariel Goldstein

Abstract

This paper investigates the problem-solving capabilities of Large Language Models (LLMs) by evaluating their performance on stumpers, unique single-step intuition problems that pose challenges for human solvers but are easily verifiable. We compare the performance of four state-of-the-art LLMs (Davinci-2, Davinci-3, GPT-3.5-Turbo, GPT-4) to human participants. Our findings reveal that the new-generation LLMs excel in solving stumpers and surpass human performance. However, humans exhibit superior skills in verifying solutions to the same problems. This research enhances our understanding of LLMs' cognitive abilities and provides insights for enhancing their problem-solving potential across various domains.

Large Language ModelsProblem-solving abilitiesStumpersCognitive abilitiesHuman performanceRiddles
BibTeX
@inproceedings{
goldstein2023decoding,
title={Decoding Stumpers: Large Language Models vs. Human Problem-Solvers},
author={Alon Goldstein and Miriam Havin and Roi Reichart and Ariel Goldstein},
booktitle={The 2023 Conference on Empirical Methods in Natural Language Processing},
year={2023},
url={https://openreview.net/forum?id=HYxJoAWLgT}
}
Decoding Stumpers: Large Language Models vs. Human Problem-Solvers · EMNLP 2023