2026
CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
ICML 2026poster
LLM-powered coding agents are reshaping the development paradigm. However, existing evaluation systems, neither traditional tests for humans nor benchmarks for LLMs, fail to capture this shift, excluding problems that require both human reasoning to guide solutions and AI efficiency for implementati…