FIXME: Towards End-to-End Benchmarking of LLM-Aided Design Verification
Gwok-Waa Wan, SamZaak Wong, Shengchu Su, Chenxu Niu, Ning Wang, Xinlai Wan, Qixiang Chen, Mengnv Xing
Abstract
We introduce FIXME, the first end-to-end and large-scale benchmark for evaluating Large Language Models (LLMs) in hardware design functional verification (FV). Comprising 747 tasks derived from real-world hardware designs, FIXME spans five core FV sub-sets: specification comprehension, reference model generation, testbench generation, assertion design, and RTL debugging. To ensure high data quality, we developed an AI-human collaborative framework for agile data curation and annotation. This process resulted in 25,000 lines of verified RTL, 35,000 lines of enhanced testbenches, and over 1,200 SystemVerilog Assertions. Furthermore, through expert-guided optimization within the multi-agent aided flow, we achieved a remarkable 45.57% improvement in average functional coverage, underscoring the benchmark
BibTeX
@inproceedings{aaai2026_fixmetowardsendt,
title = {FIXME: Towards End-to-End Benchmarking of LLM-Aided Design Verification},
author = {Gwok-Waa Wan and SamZaak Wong and Shengchu Su and Chenxu Niu and Ning Wang and Xinlai Wan and Qixiang Chen and Mengnv Xing and Jingyi Zhang and Jianmin Ye and Yubo Wang and Rongchang Song and Tao Ni and Qiang Xu and Nan Guan and Zhe Jiang and Xi Wang and Yong Chen and Jun Yang},
booktitle = {AAAI 2026},
year = {2026}
}