2026
ProbeLLM: Automating Principled Diagnosis of LLM Failures
ICML 2026poster
Understanding how and why large language models (LLMs) fail is becoming a central challenge as models rapidly evolve and static evaluations fall behind. While automated probing has been enabled by dynamic test generation, existing approaches often discover isolated failure cases, lack principled con…