← Search

Hisham Alyahya

1 accepted papers

2024

When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

ACL 2024long

Large Language Model (LLM) leaderboards based on benchmark rankings are regularly used to guide practitioners in model selection. Often, the published leaderboard rankings are taken at face value — we show this is a (potentially costly) mistake. Under existing leaderboards, the relative performance…