2026
HeurekaBench: A Benchmarking Framework for AI Co-scientist
ICLR 2026poster
LLM-based reasoning models have enabled the development of agentic systems that act as co-scientists, assisting in multi-step scientific analysis. However, evaluating these systems is challenging, as it requires realistic, end-to-end research scenarios that integrate data analysis, interpretation, a…