2026
InnoGym: Benchmarking the Innovation Potential of AI Agents
ICLR 2026poster
LLMs and Agents have achieved impressive progress in code generation, mathematical reasoning, and scientific discovery. However, existing benchmarks primarily measure correctness, overlooking the diversity of methods behind solutions. True innovation depends not only on producing correct answers but…