AAAI 2026technical0 citations

OSVBench: Benchmarking LLMs on Specification Generation Tasks for Operating System Verification

Shangyu Li, Juyong Jiang, Tiancheng Zhao, Jiasi Shen

Abstract

We introduce OSVBench, a new benchmark for evaluating Large Language Models (LLMs) on the task of generating complete formal specifications for verifying the functional correctness of operating system kernels. This benchmark is built upon a real-world operating system kernel, Hyperkernel, and consists of 245 complex specification generation tasks in total, each of which is a long-context task of about 20k-30k tokens. The benchmark formulates the specification generation task as a program synthesis problem confined to a domain for specifying states and transitions. This formulation is provided to LLMs through a programming model. The LLMs must be able to understand the programming model and verification assumptions before delineating the correct search space for syntax and semantics and generating formal specifications. Guided by the operating system

BibTeX
@inproceedings{aaai2026_osvbenchbenchmar,
  title = {OSVBench: Benchmarking LLMs on Specification Generation Tasks for Operating System Verification},
  author = {Shangyu Li and Juyong Jiang and Tiancheng Zhao and Jiasi Shen},
  booktitle = {AAAI 2026},
  year = {2026}
}
OSVBench: Benchmarking LLMs on Specification Generation Tasks for Operating System Verification · AAAI 2026