2026
XDomainBench: Diagnosing Reasoning Collapse in High-Dimensional Scientific Knowledge Composition
ICML 2026poster
Large Language Models (LLMs) are increasingly deployed for knowledge synthesis, yet their capacity for compositional generalization in scientific knowledge remains under-characterized. Existing benchmarks primarily focus on single-turn restricted scenarios, failing to capture the capability boundari…