AAAI 2025technical0 citations
Truth Behind the Scene: Designing Evaluations Benchmarks to Assess LLMs’ Task-Specific Understanding over Test-Taking Strategies
Abstract
Many existing benchmarks, such as MMLU, are limited to measuring large language models’ (LLM) true task understanding due to their reliance on statistical patterns in the training data. We suggest new approaches to improve how benchmarks can capture task-specific understanding in LLMs, revealing insights into their reasoning ability.
BibTeX
@article{Pham_2025, title={Truth Behind the Scene: Designing Evaluations Benchmarks to Assess LLMs’ Task-Specific Understanding over Test-Taking Strategies}, volume={39}, url={https://ojs.aaai.org/index.php/AAAI/article/view/35337}, DOI={10.1609/aaai.v39i28.35337}, abstractNote={Many existing benchmarks, such as MMLU, are limited to measuring large language models’ (LLM) true task understanding due to their reliance on statistical patterns in the training data. We suggest new approaches to improve how benchmarks can capture task-specific understanding in LLMs, revealing insights into their reasoning ability.}, number={28}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Pham, Thao}, year={2025}, month={Apr.}, pages={29596-29598} }