2025
Towards Optimal Evaluation Efficiency for Large Language Models
EMNLP 2025
Comprehensive evaluation of large language models (LLMs) typically requires large-scale benchmarks, which is costly in terms of both data annotation and computational resource needed for evaluation. To mitigate these challenges, we propose an efficient evaluation framework that selects a question su