ICML 2026poster0 citations

Efficient RL Training for LLMs with Experience Replay

Charles Arnal, Vivien Cabannnes, Taco Cohen, Julia Kempe, REMI MUNOS

Abstract

While Experience Replay—the practice of storing rollouts and reusing them multiple times during training—is a foundational technique in general RL, it remains largely unexplored in LLM post-training due to the prevailing belief that fresh, on-policy data is essential for high performance. In this work, we challenge this assumption. We present a systematic study of replay buffers for LLM post-training, formalizing the optimal design as a trade-off between staleness-induced variance, sample diversity and the high computational cost of generation. We show that strict on-policy sampling is suboptimal when generation is expensive. Empirically, we show that a well-designed replay buffer can drastically reduce inference compute without degrading -- and in some cases even improving -- final model performance, while preserving policy entropy.

LLMRL
BibTeX
@inproceedings{
arnal2026efficient,
title={Efficient {RL} Training for {LLM}s with Experience Replay},
author={Charles Arnal and Vivien Cabannes and Taco Cohen and Julia Kempe and R{\'e}mi Munos},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=ZlnD9QlfZh}
}