AAAI 2025technical3 citations

Real-Time Recurrent Reinforcement Learning

Julian Lemmel, Radu Grosu

Abstract

We introduce a biologically plausible RL framework for solving tasks in partially observable Markov decision processes (POMDPs). The proposed algorithm combines three integral parts: (1) A Meta-RL architecture, resembling the mammalian basal ganglia; (2) A biologically plausible reinforcement learning algorithm, exploiting temporal difference learning and eligibility traces to train the policy and the value-function; (3) An online automatic differentiation algorithm for computing the gradients with respect to parameters of a shared recurrent network backbone. Our experimental results show that the method is capable of solving a diverse set of partially observable reinforcement learning tasks. The algorithm we call real-time recurrent reinforcement learning (RTRRL) serves as a model of learning in biological neural networks, mimicking reward pathways in the basal ganglia.

BibTeX
@article{Lemmel_Grosu_2025, title={Real-Time Recurrent Reinforcement Learning}, volume={39}, url={https://ojs.aaai.org/index.php/AAAI/article/view/34001}, DOI={10.1609/aaai.v39i17.34001}, abstractNote={We introduce a biologically plausible RL framework for solving tasks in partially observable Markov decision processes (POMDPs). The proposed algorithm combines three integral parts: (1) A Meta-RL architecture, resembling the mammalian basal ganglia; (2) A biologically plausible reinforcement learning algorithm, exploiting temporal difference learning and eligibility traces to train the policy and the value-function; (3) An online automatic differentiation algorithm for computing the gradients with respect to parameters of a shared recurrent network backbone. Our experimental results show that the method is capable of solving a diverse set of partially observable reinforcement learning tasks. The algorithm we call real-time recurrent reinforcement learning (RTRRL) serves as a model of learning in biological neural networks, mimicking reward pathways in the basal ganglia.}, number={17}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Lemmel, Julian and Grosu, Radu}, year={2025}, month={Apr.}, pages={18189-18197} }
Real-Time Recurrent Reinforcement Learning · AAAI 2025