NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
Jingzhe Ding, Shengda Long, Changxin Pu, Ge Zhang, zhou huan, Hongwan Gao, Xiang Gao, Chao He
Abstract
Recent advances in coding agents suggest rapid progress toward autonomous software development, yet existing benchmarks primarily evaluate short-horizon behaviors such as localized code generation, scaffolded completion, or repository repair, leaving it unclear whether agents can sustain coherent reasoning, planning, and execution over the extended horizons demanded by real-world repository construction. To address this gap, we introduce NL2Repo-Bench, a benchmark explicitly designed to evaluate the long-horizon repository generation from scratch: given only a single natural-language requirements document and an empty workspace, agents must autonomously design the architecture, manage dependencies, and produce a fully installable Python library. Experiments across state-of-the-art open- and closed-source models reveal that long-horizon repository generation remains largely unsolved, with even the strongest agents achieving merely 40\% average test pass rates and rarely completing an entire repository correctly. Further analysis identifies systematic long-horizon failure modes, including premature termination, loss of global coherence, fragile cross-file dependencies, and inadequate planning over hundreds of interaction steps. These results position NL2Repo-Bench as a rigorous, execution-based testbed for evaluating sustained agentic competence and highlight long-horizon reasoning as a key bottleneck for autonomous coding agents. Our data and code are available at https://anonymous.4open.science/r/nl2repobench-foricml-F4ED/.
BibTeX
@inproceedings{
ding2026nlrepobench,
title={{NL}2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents},
author={Jingzhe Ding and Shengda Long and Changxin PU and Ge Zhang and zhou huan and Hongwan Gao and Xiang Gao and Chao He and Yue Hou and FEI HU and Zhaojian Li and Weiran Shi and Zaiyuan Wang and Daoguang Zan and Chenchen Zhang and Xiaoxu Zhang and Chen Qizhi and Xianfu Cheng and Bo Deng and Qingshui Gu and Kai Hua and Juntao Lin and Pai Liu and Mingchen Li and Minghao Li and Xuanguang Pan and Zifan Peng and Yujia Qin and Yong Shan and Zhewen Tan and Haoran Wang and Zihan Wang and Weihao Xie and Yishuo Yuan and Jiayu Zhang and Yunfei Zhao and He Zhu and LIYA ZHU and Chenyang Zou and Ming Ding and Jianpeng Jiao and Jiaheng Liu and Minghao Liu and Qian Liu and Chongyang Tao and Jian Yang and Tong Yang and Zhaoxiang Zhang and Xinjie Chen and Wenhao Huang},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=wqQam1muOQ}
}