2026
How Reasoning Evolves from Post-Training Data: An Empirical Study Using Chess
ICML 2026poster
We study how reasoning evolves in a language model -- from supervised fine-tuning (SFT) to reinforcement learning (RL) -- by analyzing how a set of theoretically-inspired datasets impacts language model performance in chess. We find that fine-tuning a model to directly predict the best move leads to…