2026
Beyond Benchmarks: Toward Causally Faithful Evaluation of Large Language Models
ICML 2026poster
Current LLM evaluations often conflate benchmark performance with intrinsic model capability. This is misleading, as observed outcomes arise from the entire evaluation system, including datasets, prompting methods, decoding parameters, and the software–hardware stack, rather than the model alone. Wh…