← Search

Alex Kipnis

1 accepted papers

2025

metabench - A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

ICLR 2025poster

Large Language Models (LLMs) vary in their abilities on a range of tasks. Initiatives such as the Open LLM Leaderboard aim to quantify these differences with several large benchmarks (sets of test items to which an LLM can respond either correctly or incorrectly). However, high correlations withi…

Cited by 0SourcePDFScholar