Learning User Preferences from Corrections on State Lattices
Nils Wilde, Dana Kulić, Stephen L. Smith
Abstract
Enabling a broader range of users to efficiently deploy autonomous mobile robots requires intuitive frameworks for specifying a robot's task and behaviour. We present a novel approach using learning from corrections (LfC), where a user is iteratively presented with a solution to a motion planning problem. Users might have preferences about parts of a robot's environment that are suitable for robot traffic or that should be avoided as well as preferences on the control actions a robot can take. The robot is initially unaware of these preferences; thus, we ask the user to provide a correction to the presented path. We assume that the user evaluates paths based on environment and motion features. From a sequence of corrections we learn weights for these features, which are then considered by the motion planner, resulting in future paths that better fit the user's preferences. We prove completeness of our algorithm and demonstrate its performance in simulations. Thereby, we show that the learned preferences yield good results not only for a set of training tasks but also for test tasks, as well as for different types of user behaviour.
BibTeX
@inproceedings{icra2020_learninguserpref,
title = {Learning User Preferences from Corrections on State Lattices},
author = {Nils Wilde and Dana Kulić and Stephen L. Smith},
booktitle = {ICRA 2020},
year = {2020}
}