← Search

Melissa Ailem

2 accepted papers

2024

Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks

ACL 2024long

Benchmarks have emerged as the central approach for evaluating Large Language Models (LLMs). The research community often relies on a model’s average performance across the test prompts of a benchmark to evaluate the model’s performance. This is consistent with the assumption that the test prompts w…

Cited by 11SourcePDFScholar