2025
metabench - A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
ICLR 2025poster
Large Language Models (LLMs) vary in their abilities on a range of tasks. Initiatives such as the Open LLM Leaderboard aim to quantify these differences with several large benchmarks (sets of test items to which an LLM can respond either correctly or incorrectly). However, high correlations withi…