SearcharxivSearch

arXiv subjects

Tomas J. Meijer

Publications and source records attributed to Tomas J. Meijer.

2 recordsLinked to original sources

Dual Control: On Exploration-Exploitation in Linear Systems

The term "dual control" refers to the dual objective of simultaneously balancing exploration and exploitation. Problems of this kind have been studied for nearly a century. This paper is devoted to theory and methodology relevant for optimal control of linear time-invariant systems whose parameters are initially unknown and must be learned by active probing. We review the main ideas underlying four major research directions: Multi-armed bandits, self-tuning regulators, regret rate minimizing controllers, and minimax optimal dual controllers. The first three have a long history and rich literature, whereas the fourth provides a promising framework for robust dual control.

math.OC

Minimax optimal dual control of positive systems: an exact solution for scalar input-sign uncertainty

While recent advances in minimax dual control have led to exact solutions for uncertain general linear time-invariant systems as well as (sub)optimal dual controllers, corresponding results for linear positive systems are still lacking. This paper aims to fill this gap and thereby pave the way toward scalable dual control algorithms. We study the general minimax optimal dual control problem for positive linear systems with unknown dynamics and reformulate it as a standard zero-sum dynamic game. By allowing randomized control inputs, we solve the corresponding Bellman equation exactly for the scalar case with sign uncertainty in the input. This yields an implicit dual control policy that is optimal both in terms of cost and $\ell_1$-gain. The optimal dual policy uses exploration in a specific region of the hyperstate space to conduct optimal probing. Outside this exploration regime, the controller reduces to a deterministic certainty equivalence policy, indicating that sufficient information has been obtained to identify the correct input direction. In addition, these results allow us to analyze fundamental limitations of minimax dual control for positive systems and provide a foundation for more general dual control problems for positive systems for future work.

math.OC