Robust exploration in linear quadratic reinforcement learning
Jack Umenberger, Mina Ferizbegovic, Thomas B Schön, Håkan Hjalmarsson
Abstract
Learning to make decisions in an uncertain and dynamic environment is a task of fundamental performance in a number of domains. This paper concerns the problem of learning control policies for an unknown linear dynamical system so as to minimize a quadratic cost function. We present a method, based on convex optimization, that accomplishes this task ‘robustly’, i.e., the worst-case cost, accounting for system uncertainty given the observed data, is minimized. The method balances exploitation and exploration, exciting the system in such a way so as to reduce uncertainty in the model parameters to which the worst-case cost is most sensitive. Numerical simulations and application to a hardware-in-the-loop servo-mechanism are used to demonstrate the approach, with appreciable performance and robustness gains over alternative methods observed in both.
BibTeX
@inproceedings{NEURIPS2019_060fd70a,
author = {Umenberger, Jack and Ferizbegovic, Mina and Sch\"{o}n, Thomas B and Hjalmarsson, H\aa kan},
booktitle = {Advances in Neural Information Processing Systems},
editor = {H. Wallach and H. Larochelle and A. Beygelzimer and F. d\textquotesingle Alch\'{e}-Buc and E. Fox and R. Garnett},
pages = {},
publisher = {Curran Associates, Inc.},
title = {Robust exploration in linear quadratic reinforcement learning},
url = {https://proceedings.neurips.cc/paper_files/paper/2019/file/060fd70a06ead2e1079d27612b84aff4-Paper.pdf},
volume = {32},
year = {2019}
}