From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old Ones
Does reinforcement learning (RL) teach large language models (LLMs) genuinely new skills, or does it merely activate existing ones? This question lies at the core of ongoing debates about the role of RL in LLM post-training. On one side, strong empirical results can be achieved with RL alone even wi…