IDE: A Multi-Agent-Driven Iterative Framework for Dynamic Evaluation of LLMs
Xin Tong, Bo Jin, Jingya Wang, Wenpeng Xing, Tian Xia, Meng Han
Abstract
With the widespread use of large language models (LLMs) in natural language processing, traditional evaluation methods based on static datasets have become inadequate to fully capture their performance and generalization capabilities. To address this challenge, we propose an Iterative Dynamic Evaluation (IDE) framework, which utilizes a multi-agent system to systematically evaluate LLMs. The innovation of IDE lies in its iterative enhancement and comparative selection mechanisms. Through successive rounds of data augmentation, the framework emulates evolutionary processes by progressively increasing sample complexity, while competitive selection ensures only the optimal samples are retained. This iterative optimization and filtering produce increasingly challenging and diverse datasets. Experimental results demonstrate that IDE more effectively exposes the limitations of models in complex tasks compared to traditional methods, providing a more comprehensive foundation for the evaluation, optimization, and application of LLMs.
BibTeX
@inproceedings{icassp2025_ideamultiagentdr,
title = {IDE: A Multi-Agent-Driven Iterative Framework for Dynamic Evaluation of LLMs},
author = {Xin Tong and Bo Jin and Jingya Wang and Wenpeng Xing and Tian Xia and Meng Han},
booktitle = {ICASSP 2025},
year = {2025}
}