← Search

Wanling Gao

2 accepted papers

2026

Beyond Benchmarks: Toward Causally Faithful Evaluation of Large Language Models

ICML 2026poster

Current LLM evaluations often conflate benchmark performance with intrinsic model capability. This is misleading, as observed outcomes arise from the entire evaluation system, including datasets, prompting methods, decoding parameters, and the software–hardware stack, rather than the model alone. Wh…

Cited by 0SourceScholar
2026

CombinationTS: A Modular Framework for Understanding Time-Series Forecasting Models

ICML 2026poster

Recent progress in time-series forecasting has led to rapidly increasing architectural complexity, yet many reported State-of-the-Art gains are statistically fragile or misattributed. We argue that progress requires a shift from model selection to modular attribution, identifying which components tr…

Cited by 0SourceScholar