← Search

Kefan Song

2 accepted papers

2026

Reward Is Enough: LLMs Are In-Context Reinforcement Learners

ICLR 2026poster

Reinforcement learning (RL) is a human-designed framework for solving sequential decision-making problems. In this work, we demonstrate that, surprisingly, RL emerges in LLMs at inference time – a phenomenon known as in-context RL (ICRL). To reveal this capability, we introduce a simple multi-round…

Cited by 0SourceScholar