← Search

Bhavya Bhavya

2 accepted papers

2025

ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

ICML 2025oral

Realizing the vision of using AI agents to automate critical IT tasks depends on the ability to measure and understand effectiveness of proposed solutions. We introduce ITBench, a framework that offers a systematic methodology for benchmarking AI agents to address real-world IT automation tasks. Our…

2025

STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds

NeurIPS 2025poster

In cloud-scale systems, failures are the norm. A distributed computing cluster exhibits hundreds of machine failures and thousands of disk failures; software bugs and misconfigurations are reported to be more frequent. The demand for autonomous, AI-driven reliability engineering continues to grow, a…

Cited by 0SourceScholar