← Search

Tianyin Xu

3 accepted papers

2026

SysMoBench: Evaluating AI on Formally Specifying Complex Real-World Systems

ICLR 2026poster

Formal models are essential to specifying large, complex computer systems and verifying their correctness, but are notoriously expensive to write and maintain. Recent advances in generative AI show promise in generating certain forms of specifications. However, existing work mostly targets small cod…

Cited by 0SourcecodeScholar
2025

ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

ICML 2025oral

Realizing the vision of using AI agents to automate critical IT tasks depends on the ability to measure and understand effectiveness of proposed solutions. We introduce ITBench, a framework that offers a systematic methodology for benchmarking AI agents to address real-world IT automation tasks. Our…

2025

STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds

NeurIPS 2025poster

In cloud-scale systems, failures are the norm. A distributed computing cluster exhibits hundreds of machine failures and thousands of disk failures; software bugs and misconfigurations are reported to be more frequent. The demand for autonomous, AI-driven reliability engineering continues to grow, a…

Cited by 0SourceScholar