2024
Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks
ACL 2024long
Benchmarks have emerged as the central approach for evaluating Large Language Models (LLMs). The research community often relies on a model’s average performance across the test prompts of a benchmark to evaluate the model’s performance. This is consistent with the assumption that the test prompts w…