SearcharxivSearch

arXiv subjects

Mathieu Granzotto

Publications and source records attributed to Mathieu Granzotto.

12 recordsLinked to original sources

On robustness, input-to-state stability and backstepping for stochastic differential equations

We study conditions under which stability of the origin of stochastic differential equations is robust to small perturbations. We express robustness in two ways, firstly in the sense that stochastic stability is maintained under small parametric perturbations not exceeding a state-dependent bound vanishing at the origin but positive elsewhere, and secondly via stochastic input-to-state stability (ISS) which allows non-zero perturbations everywhere. We prove the former property assuming the existence of a Lyapunov function certifying stochastic stability of the nominal system. Under the same assumption, stochastic ISS holds under a suitable state-dependent perturbation scaling. Stochastic exponential stability is maintained under proportionally bounded perturbations and implies exponential ISS even without perturbation scaling. Finally, we propose a novel approach to stochastic integrator backstepping in pure-feedback form that uses the tools from our robustness analysis.

eess.SY

Value iteration with stopping criterion: finite iterations, stability, and near-optimality guarantees

Value iteration (VI) is a cornerstone of dynamic programming that allows computing near-optimal feedback laws for general plant dynamics and cost functions. In practice, however, it must be stopped after finitely many iterations. This raises the question of when to stop the algorithm so that the resulting policies and value functions achieve desirable properties, like given near-optimality bounds and stability. In this context, we study deterministic, discrete-time systems with infinite-horizon (possibly discounted) costs whose inputs are generated by VI. We equip VI with a generalized stopping criterion that encompasses existing choices while allowing new ones. Our aim is to analyze the properties of the policies and value functions at the final iteration. Under mild assumptions, we first show that VI indeed terminates in a finite number of iterations. We then establish that the final policies are stabilizing by properly designing the stopping criterion, and derive explicit near-optimality bounds characterized by this choice. These results offer a design framework for the stopping criteria that balances computational effort with stability and performance guarantees.

math.OC

Discounted MPC and infinite-horizon optimal control under plant-model mismatch: Stability and suboptimality

We study closed-loop stability and suboptimality for MPC and infinite-horizon optimal control solved using a surrogate model that differs from the real plant. We employ a unified framework based on quadratic costs to analyze both finite- and infinite-horizon problems, encompassing discounted and undiscounted scenarios alike. Plant-model mismatch bounds proportional to states and controls are assumed, under which the origin remains an equilibrium. Under continuity of the model and cost-controllability, exponential stability of the closed loop can be guaranteed. Furthermore, we give a suboptimality bound for the closed-loop cost recovering the optimal cost of the surrogate. The results reveal a tradeoff between horizon length, discounting and plant-model mismatch. The robustness guarantees are uniform over the horizon length, meaning that larger horizons do not require successively smaller plant-model mismatch.

math.OC

Discounted LQR: stabilizing (near-)optimal state-feedback laws

We study deterministic, discrete linear time-invariant systems with infinite-horizon discounted quadratic cost. It is well-known that standard stabilizability and detectability properties are not enough in general to conclude stability properties for the system in closed-loop with the optimal controller when the discount factor is small. In this context, we first review some of the stability conditions based on the optimal value function found in the learning and control literature and highlight their conservatism. We then propose novel (necessary and) sufficient conditions, still based on the optimal value function, under which stability of the origin for the optimal closed-loop system is guaranteed. Afterwards, we focus on the scenario where the optimal feedback law is not stabilizing because of the discount factor and the goal is to design an alternative stabilizing near-optimal static state-feedback law. We present both linear matrix inequality-based conditions and a variant of policy iteration to construct such stabilizing near-optimal controllers. The methods are illustrated via numerical examples.

math.OC

An optimistic planning algorithm for switched discrete-time LQR

We introduce TROOP, a tree-based Riccati optimistic online planner, that is designed to generate near-optimal control laws for discrete-time switched linear systems with switched quadratic costs. The key challenge that we address is balancing computational resources against control performance, which is important as constructing near-optimal inputs often requires substantial amount of computations. TROOP addresses this trade-off by adopting an online best-first search strategy inspired by A*, allowing for efficient estimates of the optimal value function. The control laws obtained guarantee both near-optimality and stability properties for the closed-loop system. These properties depend on the planning depth, which determines how far into the future the algorithm explores and is closely related to the amount of computations. TROOP thus strikes a balance between computational efficiency and control performance, which is illustrated by numerical simulations on an example.

math.OC

Robust Recurrence of Discrete-Time Infinite-Horizon Stochastic Optimal Control with Discounted Cost

We analyze the stability of general nonlinear discrete-time stochastic systems controlled by optimal inputs that minimize an infinite-horizon discounted cost. Under a novel stochastic formulation of cost-controllability and detectability assumptions inspired by the related literature on deterministic systems, we prove that uniform semi-global practical recurrence holds for the closed-loop system, where the adjustable parameter is the discount factor. Under additional continuity assumptions, we further prove that this property is robust.

math.OC

Policy iteration for discrete-time systems with discounted costs: stability and near-optimality guarantees

Given a discounted cost, we study deterministic discrete-time systems whose inputs are generated by policy iteration (PI). We provide novel near-optimality and stability properties, while allowing for non stabilizing initial policies. That is, we first give novel bounds on the mismatch between the value function generated by PI and the optimal value function, which are less conservative in general than those encountered in the dynamic programming literature for the considered class of systems. Then, we show that the system in closed-loop with policies generated by PI are stabilizing under mild conditions, after a finite (and known) number of iterations.

math.OC

Policy iteration: for want of recursive feasibility, all is not lost

This paper investigates recursive feasibility, recursive robust stability and near-optimality properties of policy iteration (PI). For this purpose, we consider deterministic nonlinear discrete-time systems whose inputs are generated by PI for undiscounted cost functions. We first assume that PI is recursively feasible, in the sense that the optimization problems solved at each iteration admit a solution. In this case, we provide novel conditions to establish recursive robust stability properties for a general attractor, meaning that the policies generated at each iteration ensure a robust \KL-stability property with respect to a general state measure. We then derive novel explicit bounds on the mismatch between the (suboptimal) value function returned by PI at each iteration and the optimal one. Afterwards, motivated by a counter-example that shows that PI may fail to be recursively feasible, we modify PI so that recursive feasibility is guaranteed a priori under mild conditions. This modified algorithm, called PI+, is shown to preserve the recursive robust stability when the attractor is compact. Additionally, PI+ enjoys the same near-optimality properties as its PI counterpart under the same assumptions. Therefore, PI+ is an attractive tool for generating near-optimal stabilizing control of deterministic discrete-time nonlinear systems.

math.OC

Stability analysis of optimal control problems with time-dependent costs

We present stability conditions for deterministic time-varying nonlinear discrete-time systems whose inputs aim to minimize an infinite-horizon time-dependent cost. Global asymptotic and exponential stability properties for general attractors are established. This work covers and generalizes the related results on discounted optimal control problems to more general systems and cost functions.

eess.SY

Exploiting homogeneity for the optimal control of discrete-time systems: application to value iteration

To investigate solutions of (near-)optimal control problems, we extend and exploit a notion of homogeneity recently proposed in the literature for discrete-time systems. Assuming the plant dynamics is homogeneous, we first derive a scaling property of its solutions along rays provided the sequence of inputs is suitably modified. We then consider homogeneous cost functions and reveal how the optimal value function scales along rays. This result can be used to construct (near-)optimal inputs on the whole state space by only solving the original problem on a given compact manifold of a smaller dimension. Compared to the related works of the literature, we impose no conditions on the homogeneity degrees. We demonstrate the strength of this new result by presenting a new approximate scheme for value iteration, which is one of the pillars of dynamic programming. The new algorithm provides guaranteed lower and upper estimates of the true value function at any iteration and has several appealing features in terms of reduced computation. A numerical case study is provided to illustrate the proposed algorithm.

math.OC

When to stop value iteration: stability and near-optimality versus computation

Value iteration (VI) is a ubiquitous algorithm for optimal control, planning, and reinforcement learning schemes. Under the right assumptions, VI is a vital tool to generate inputs with desirable properties for the controlled system, like optimality and Lyapunov stability. As VI usually requires an infinite number of iterations to solve general nonlinear optimal control problems, a key question is when to terminate the algorithm to produce a "good" solution, with a measurable impact on optimality and stability guarantees. By carefully analysing VI under general stabilizability and detectability properties, we provide explicit and novel relationships of the stopping criterion's impact on near-optimality, stability and performance, thus allowing to tune these desirable properties against the induced computational cost. The considered class of stopping criteria encompasses those encountered in the control, dynamic programming and reinforcement learning literature and it allows considering new ones, which may be useful to further reduce the computational cost while endowing and satisfying stability and near-optimality properties. We therefore lay a foundation to endow machine learning schemes based on VI with stability and performance guarantees, while reducing computational complexity.

math.OC

Optimistic planning for the near-optimal control of nonlinear switched discrete-time systems with stability guarantees

Originating in the artificial intelligence literature, optimistic planning (OP) is an algorithm that generates near-optimal control inputs for generic nonlinear discrete-time systems whose input set is finite. This technique is therefore relevant for the near-optimal control of nonlinear switched systems, for which the switching signal is the control. However, OP exhibits several limitations, which prevent its application in a standard control context. First, it requires the stage cost to take values in [0,1], an unnatural prerequisite as it excludes, for instance, quadratic stage costs. Second, it requires the cost function to be discounted. Third, it applies for reward maximization, and not cost minimization. In this paper, we modify OP to overcome these limitations, and we call the new algorithm OPmin. We then make stabilizability and detectability assumptions, under which we derive near-optimality guarantees for OPmin and we show that the obtained bound has major advantages compared to the bound originally given by OP. In addition, we prove that a system whose inputs are generated by OPmin in a receding-horizon fashion exhibits stability properties. As a result, OPmin provides a new tool for the near-optimal, stable control of nonlinear switched discrete-time systems for generic cost functions.

math.OC