2026
CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments
ICML 2026poster
Current evaluation for Large Language Model (LLM) code agents predominantly focus on generating functional code in single-turn scenarios, which fails to evaluate the agent's capability for continuous code optimization and multi-turn iterative development. To bridge this gap, we introduce CATArena, a…