LLM4Cov: Execution-Grounded Agent Learning for High-Coverage Hardware Verification
Hejia Zhang, Zhongming Yu, Chia-Tung Ho, Mark Ren, Brucek Khailany, Jishen Zhao
Abstract
Execution-grounded LLM agents offer a promising paradigm for learning from tool feedback, but such feedback is often expensive and slow to obtain, making online reinforcement learning (RL) impractical. High-coverage hardware verification exemplifies this challenge due to its reliance on industrial simulators and non-differentiable execution signals. We propose LLM4Cov, an offline agent-learning framework that models verification as memoryless state transitions guided by deterministic evaluators. Building on this formulation, we introduce execution-validated data curation, policy-aware agentic data synthesis, and worst-state-prioritized sampling to enable scalable learning under execution constraints. We further curate a reality-aligned benchmark adapted from an existing verification suite through a revised evaluation protocol. Using the proposed pipeline, a compact 4B-parameter model achieves 69.2\% coverage pass rate under agentic evaluation, outperforming its teacher by 5.3\% and demonstrating competitive performance against models an order of magnitude larger.
BibTeX
@inproceedings{
zhang2026llmcov,
title={{LLM}4Cov: Execution-Aware Agentic Learning for High-coverage Testbench Generation},
author={Hejia Zhang and Zhongming Yu and Chia-Tung Ho and Haoxing Ren and Brucek Khailany and Jishen Zhao},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=q8AIRg06GX}
}