2026
Don’t Pass@$k$: A Bayesian Framework for Large Language Model Evaluation
ICLR 2026poster
Pass@$k$ is widely used to report performance for LLM reasoning, but it often yields unstable, misleading rankings, especially when the number of trials (samples) is limited and compute is constrained. We present a principled Bayesian evaluation framework that replaces Pass@$k$ and average accuracy…