arXiv · 1906.01584
Robust exploration in linear quadratic reinforcement learning
Abstract
This paper concerns the problem of learning control policies for an unknown linear dynamical system to minimize a quadratic cost function. We present a method, based on convex optimization, that accomplishes this task robustly: i.e., we minimize the worst-case cost, accounting for system uncertainty given the observed data. The method balances exploitation and exploration, exciting the system in such a way so as to reduce uncertainty in the model parameters to which the worst-case cost is most sensitive. Numerical simulations and application to a hardware-in-the-loop servo-mechanism demonstrate the approach, with appreciable performance and robustness gains over alternative methods observed in both.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jack Umenberger, Mina Ferizbegovic, Thomas B. Schön, Håkan Hjalmarsson. 2019-06-04. Robust exploration in linear quadratic reinforcement learning. https://arxiv.org/abs/1906.01584
Cite the original work for its findings. Save a collection to share your selection of sources.