← Search

Mudit Verma

5 accepted papers

2025

ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

ICML 2025oral

Realizing the vision of using AI agents to automate critical IT tasks depends on the ability to measure and understand effectiveness of proposed solutions. We introduce ITBench, a framework that offers a systematic methodology for benchmarking AI agents to address real-world IT automation tasks. Our…

2024

Position: LLMs Can’t Plan, But Can Help Planning in LLM-Modulo Frameworks

ICML 2024spotlight

We argue that auto-regressive LLMs cannot, by themselves, do planning or self-verification (which is after all a form of reasoning), and shed some light on the reasons for misunderstandings in the literature. We will also argue that LLMs should be viewed as universal approximate knowledge sources th…

Cited by 192SourcePDFScholar
2022

Bridging the Gap: Providing Post-Hoc Symbolic Explanations for Sequential Decision-Making Problems with Inscrutable Representations

ICLR 2022poster

As increasingly complex AI systems are introduced into our daily lives, it becomes important for such systems to be capable of explaining the rationale for their decisions and allowing users to contest these decisions. A significant hurdle to allowing for such explanatory dialogue could be the {\em…

Cited by 42SourcePDFScholar
2021

Widening the Pipeline in Human-Guided Reinforcement Learning with Explanation and Context-Aware Data Augmentation

NeurIPS 2021spotlight

Human explanation (e.g., in terms of feature importance) has been recently used to extend the communication channel between human and agent in interactive machine learning. Under this setting, human trainers provide not only the ground truth but also some form of explanation. However, this kind of h…

Cited by 50SourcePDFScholar