2025
ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
EMNLP 2025
LLMs are highly sensitive to prompt phrasing, yet standard benchmarks typically report performance using a single prompt, raising concerns about the reliability of such evaluations. In this work, we argue for a stochastic method of moments evaluation over the space of meaning-preserving prompt pertu