ICLR 2026poster0 citations

SysMoBench: Evaluating AI on Formally Specifying Complex Real-World Systems

Qian Cheng, Ruize Tang, Emilie Ma, Finn Hackett, Peiyang He, Yiming Su, Ivan Beschastnikh, Yu Huang

Abstract

Formal models are essential to specifying large, complex computer systems and verifying their correctness, but are notoriously expensive to write and maintain. Recent advances in generative AI show promise in generating certain forms of specifications. However, existing work mostly targets small code, not complete systems. It is unclear whether AI can deal with realistic system artifacts, as this requires abstracting their complex behavioral properties into formal models. We present SysMoBench, a benchmark that evaluates AI's ability to formally model large, complex systems. We focus on concurrent and distributed systems, which are keystones of today's critical computing infrastructures, encompassing operating systems and cloud infrastructure. We use TLA+, the de facto specification language for concurrent and distributed systems, though the benchmark can be extended to other specification languages. We address the primary challenge of evaluating AI-generated models by automating metrics like syntactic and runtime correctness, conformance to system code, and invariant correctness. SysMoBench currently includes eleven diverse system artifacts: the Raft implementation of Etcd and Redis, the leader election of ZooKeeper, the Spinlock, Mutex and Ringbuffer in Asterinas OS, etc., with more being added. SysMoBench enables us to understand the capabilities and limitations of today's LLMs and agents, putting tools in this area on a firm footing and opening up promising new research directions.

SpecificationBenchmarkDistributed SystemConcurrent SystemAgentic AILarge Language Model
BibTeX
@inproceedings{
cheng2026sysmobench,
title={SysMoBench: Evaluating {AI} on Formally Specifying Complex Real-World Systems},
author={Qian Cheng and Ruize Tang and Emilie Ma and Finn Hackett and Peiyang He and Yiming Su and Ivan Beschastnikh and Yu Huang and Xiaoxing Ma and Tianyin Xu},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=SAeaTz8YoM}
}
SysMoBench: Evaluating AI on Formally Specifying Complex Real-World Systems · ICLR 2026