← Search

Yunyou Huang

1 accepted papers

2026

Beyond Benchmarks: Toward Causally Faithful Evaluation of Large Language Models

ICML 2026poster

Current LLM evaluations often conflate benchmark performance with intrinsic model capability. This is misleading, as observed outcomes arise from the entire evaluation system, including datasets, prompting methods, decoding parameters, and the software–hardware stack, rather than the model alone. Wh…

Cited by 0SourceScholar