2026
SEDRAS: Symbolically Evaluated Deep Research And Science
ICML 2026poster
As the reasoning capabilities of Large Language Models (LLMs) expand, evaluating true inductive generalization on entirely unseen data becomes increasingly challenging. To this end, we introduce a modular in-context learning evaluation framework, that is scalable and extendable across its separate m…