2026
Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
ICLR 2026poster
The paradigm shift in large language models (LLMs) from instinctive responses to chain-of-thought (CoT) reasoning has fueled two prevailing assumptions: (1) reasoning capabilities only emerge in sufficiently large models, and (2) such capabilities require training on massive datasets. While the firs…