← Search

Aafiya Shamshad Hussain

1 accepted papers

2026

BeyondBench: Benchmark-Free Evaluation of Reasoning in Language Models

ICLR 2026poster

Evaluating language models fairly is becoming harder as static benchmarks available on the internet risk contamination by training data. This makes it unclear whether models are truly reasoning or just recalling answers. In this paper, we introduce $\textbf{BeyondBench}$, an evaluation framework tha…

Cited by 0SourcecodeScholar