2026
D-FUSEr: Diverse Failure, Unified Success via Error-Distribution Shaping in LLM Reasoning
ICML 2026poster
Test-time scaling methods such as majority vote aggregation and iterative refinement (e.g., self-reflection or multi-agent inference) improve reasoning performance by leveraging multiple solution samples. However, their efficacy depends not only on raw performance, but critically on the distribution…