2026
FIRE-Bench: Evaluating Agents on the Rediscovery of Scientific Insights
ICML 2026poster
Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery, but rigorously evaluating their capacity for verifiable discovery remains a central challenge. Existing benchmarks face a trade-off: they either rely on LLM-as-judge evaluations of automatically gen…