SearcharxivSearch

arXiv subjects

Simone Baroncini

Publications and source records attributed to Simone Baroncini.

2 recordsLinked to original sources

On Reward-Balancing Methods for Reinforcement Learning

This paper investigates the so-called reward-balancing methods, a novel class of algorithms for solving discounted-return reinforcement learning (RL) problems. These methods consist of iteratively adjusting the reward function to transform the RL problem into an equivalent one in which the optimal policies are greedy. For this procedure, referred to as normalization process, we provide a theoretical analysis of the involved transformations, emphasizing their algebraic structure. Then, we introduce a control-theoretic reformulation, recasting the reward-balancing procedure into an optimal control framework. The approach is further extended to address model uncertainty through stochastic model sampling, yielding normalization guarantees and probabilistic bounds on stochastic fluctuations. Using the proposed optimal control framework within a scenario model predictive control (MPC) setting, we demonstrate, through simulation studies, performance improvements over the current state-of-the-art.

math.OC

On Sufficient Richness for Linear Time-Invariant Systems

Persistent excitation (PE) is a necessary and sufficient condition for uniform exponential parameter convergence in several adaptive, identification, and learning schemes. In this article, we consider, in the context of multi-input linear time-invariant (LTI) systems, the problem of guaranteeing PE of commonly-used regressors by applying a sufficiently rich (SR) input signal. Exploiting the analogies between time shifts and time derivatives, we state simple necessary and sufficient PE conditions for the discrete- and continuous-time frameworks. Moreover, we characterize the shape of the set of SR input signals for both single-input and multi-input systems. Finally, we show with a numerical example that the derived conditions are tight and cannot be improved without including additional knowledge of the considered LTI system.

eess.SY