2026
Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently
ICML 2026poster
Transformers can acquire Chain-of-Thought (CoT) capabilities to solve complex reasoning tasks through fine-tuning. Reinforcement learning (RL) and supervised fine-tuning (SFT) are two primary approaches to this end. In this work, we examine them specifically for learning k-sparse Boolean functions w…