2026
AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
ICML 2026poster
As LLM-based agents are increasingly deployed in real-world workflows, existing agent benchmarks---often built on idealized, noise-free assumptions---fall short of characterizing agents' robustness under imperfect user instructions and unreliable tool feedback. To address this gap, we introduce **Ag…