ICLR 2021poster26 citations

Balancing Constraints and Rewards with Meta-Gradient D4PG

Dan A. Calian, Daniel J Mankowitz, Tom Zahavy, Zhongwen Xu, Junhyuk Oh, Nir Levine, Timothy Mann

Abstract

Deploying Reinforcement Learning (RL) agents to solve real-world applications often requires satisfying complex system constraints. Often the constraint thresholds are incorrectly set due to the complex nature of a system or the inability to verify the thresholds offline (e.g, no simulator or reasonable offline evaluation procedure exists). This results in solutions where a task cannot be solved without violating the constraints. However, in many real-world cases, constraint violations are undesirable yet they are not catastrophic, motivating the need for soft-constrained RL approaches. We present two soft-constrained RL approaches that utilize meta-gradients to find a good trade-off between expected return and minimizing constraint violations. We demonstrate the effectiveness of these approaches by showing that they consistently outperform the baselines across four different Mujoco domains.

reinforcement learningmeta-gradientsconstraints
BibTeX
@inproceedings{
calian2021balancing,
title={Balancing Constraints and Rewards with Meta-Gradient D4{\{}PG{\}}},
author={Dan A. Calian and Daniel J Mankowitz and Tom Zahavy and Zhongwen Xu and Junhyuk Oh and Nir Levine and Timothy Mann},
booktitle={International Conference on Learning Representations},
year={2021},
url={https://openreview.net/forum?id=TQt98Ya7UMP}
}
Balancing Constraints and Rewards with Meta-Gradient D4PG · ICLR 2021