2026
Toward More Reliable Agent Evaluation: A Component-Based Benchmark Auditing Pipeline
ICML 2026poster
Reliable evaluation of large language model (LLM) agents depends critically on benchmark validity. However, agent benchmarks are increasingly complex and often contain hidden flaws arising from interactions among user instructions, environments, tools, ground-truth trajectories, and evaluation proto…