ICLR 2026poster0 citations

Towards Self-Evolving Agent Benchmarks : Validatable Agent Trajectory via Test-Time Exploration

Dadi Guo, Tianyi Zhou, Dongrui Liu, Chen Qian, Qihan Ren, Shuai Shao, Zhiyuan Fan, Yi R. Fung

Abstract

Recent advances in large language models (LLMs) and agent system designs have empowered agents with unprecedented levels of capability. However, existing agent benchmarks are showing a trend of rapid ceiling-hitting by newly developed agents, making it difficult to meet the demands for evaluating agent abilities. To address this problem, we propose the Trajectory-based Reproducible Agent- benchmark Complexity Evolution (TRACE) framework. This framework takes an original task from an existing benchmark and encourages agents to freely explore and evolve it into a new task with higher difficulty while recording traceable agent trajectories. The framework proceeds in three stages: (1) evolutionary proposal mining, which provides task evolution proposals through preliminary exploration and divergent thinking; (2) problem formation and free exploration, where proposals are conceptualized into feasible problem candidates and the agents then explore them freely while recording their execution trajectories; and (3) multi-level validation, which ensures that the evolved tasks are accompanied by validatable and reproducible trajectories. Experiments on the GAIA benchmark demonstrate that the TRACE framework consistently enhances task complexity while improving the reliability of correctness through validatable execution trajectories. This work marks a paradigm shift from static, manually curated benchmarks to dynamic, self-evolving evaluation systems, providing a sustainable and challenging runway for agent development

Benchmark EvolutionAgent EvaluationTest-Time ExplorationMulti-Agent SystemsLarge Language ModelsDynamic Task Generation
BibTeX
@inproceedings{
guo2026towards,
title={Towards Self-Evolving Agent Benchmarks : Validatable Agent Trajectory via Test-Time Exploration},
author={Dadi Guo and Tianyi Zhou and Dongrui Liu and Chen Qian and Qihan Ren and Shuai Shao and Zhiyuan Fan and Yi R. Fung and Kun Wang and Linfeng Zhang and Jing Shao},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=2H03gm4Rq6}
}