AAAI 2025technical0 citations

On Shallow Planning Under Partial Observability

Randy Lefebvre, Audrey Durand

Abstract

Formulating a real-world problem under the Reinforcement Learning framework involves non-trivial design choices, such as selecting a discount factor for the learning objective (dis- counted cumulative rewards), which articulates the planning horizon of the agent. This work investigates the impact of the discount factor on the bias-variance trade-off given structural parameters of the underlying Markov Decision Process. Our results support the idea that a shorter planning horizon might be beneficial, especially under partial observability.

BibTeX
@article{Lefebvre_Durand_2025, title={On Shallow Planning Under Partial Observability}, volume={39}, url={https://ojs.aaai.org/index.php/AAAI/article/view/34860}, DOI={10.1609/aaai.v39i25.34860}, abstractNote={Formulating a real-world problem under the Reinforcement
Learning framework involves non-trivial design choices, such
as selecting a discount factor for the learning objective (dis-
counted cumulative rewards), which articulates the planning
horizon of the agent. This work investigates the impact of the
discount factor on the bias-variance trade-off given structural
parameters of the underlying Markov Decision Process. Our
results support the idea that a shorter planning horizon might
be beneficial, especially under partial observability.}, number={25}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Lefebvre, Randy and Durand, Audrey}, year={2025}, month={Apr.}, pages={26587-26595} }