AAAI 2025technical0 citations

Evaluating LLM Reasoning in the Operations Research Domain with ORQA

Mahdi Mostajabdaveh, Timothy Tin Long Yu, Samarendra Chandan Bindu Dash, Rindra Ramamonjison, Jabo Serge Byusa, Giuseppe Carenini, Zirui Zhou, Yong Zhang

Abstract

In this paper, we introduce and apply Operations Research Question Answering (ORQA), a new benchmark, to assess the generalization capabilities of Large Language Models (LLMs) in the specialized technical domain of Operations Research (OR). This benchmark is designed to evaluate whether LLMs can emulate the knowledge and reasoning skills of OR experts when given diverse and complex optimization problems. The dataset, crafted by OR experts, presents real-world optimization problems that require multistep reasoning to build their mathematical models. Our evaluations of various open-source LLMs, such as LLaMA 3.1, DeepSeek, and Mixtral reveal their modest performance, indicating a gap in their aptitude to generalize to specialized technical domains. This work contributes to the ongoing discourse on LLMs’ generalization capabilities, providing insights for future research in this area. The dataset and evaluation code are publicly available.

BibTeX
@article{Mostajabdaveh_Yu_Dash_Ramamonjison_Byusa_Carenini_Zhou_Zhang_2025, title={Evaluating LLM Reasoning in the Operations Research Domain with ORQA}, volume={39}, url={https://ojs.aaai.org/index.php/AAAI/article/view/34673}, DOI={10.1609/aaai.v39i23.34673}, abstractNote={In this paper, we introduce and apply Operations Research Question Answering (ORQA), a new benchmark, to assess the generalization capabilities of Large Language Models (LLMs) in the specialized technical domain of Operations Research (OR). This benchmark is designed to evaluate whether LLMs can emulate the knowledge and reasoning skills of OR experts when given diverse and complex optimization problems. The dataset, crafted by OR experts, presents real-world optimization problems that require multistep reasoning to build their mathematical models. Our evaluations of various open-source LLMs, such as LLaMA 3.1, DeepSeek, and Mixtral reveal their modest performance, indicating a gap in their aptitude to generalize to specialized technical domains. This work contributes to the ongoing discourse on LLMs’ generalization capabilities, providing insights for future research in this area. The dataset and evaluation code are publicly available.}, number={23}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Mostajabdaveh, Mahdi and Yu, Timothy Tin Long and Dash, Samarendra Chandan Bindu and Ramamonjison, Rindra and Byusa, Jabo Serge and Carenini, Giuseppe and Zhou, Zirui and Zhang, Yong}, year={2025}, month={Apr.}, pages={24902-24910} }
Evaluating LLM Reasoning in the Operations Research Domain with ORQA · AAAI 2025