← Search

Amirhossein Samandar

1 accepted papers

2026

Don’t Pass@$k$: A Bayesian Framework for Large Language Model Evaluation

ICLR 2026poster

Pass@$k$ is widely used to report performance for LLM reasoning, but it often yields unstable, misleading rankings, especially when the number of trials (samples) is limited and compute is constrained. We present a principled Bayesian evaluation framework that replaces Pass@$k$ and average accuracy…

Cited by 0SourcecodeScholar