Searcharxiv⌕ Search

arXiv subjects

Michele Palladino

Publications and source records attributed to Michele Palladino.

15 recordsLinked to original sources

Minimizers that are not Impulsive Minimizers and Higher Order Abnormality

This paper addresses two related problems in optimal control. The first investigation consists of compatibility issues between two classical approaches to deriving necessary conditions for optimal control problems with a final target: the set-separation approach and penalization techniques. These methods generally lead to non-equivalent conditions, mainly due to their reliance on different notions of tangency at the target. We address this issue by considering Quasi Differential Quotient (QDQ) approximating cones (which are fit for the set-separation approach) and identifying conditions under which the Clarke tangent cone (which is a typical tool within penalization techniques) is also a QDQ approximating cone. In particular, we show that this property holds under suitable local invariance assumptions or when the target coincides locally with an $r$-prox regular set. In the second part of the paper we apply this compatibility result to the study of infimum-gap phenomena in optimal control problems with unbounded controls and impulsive extensions. In particular, we establish a connection between the occurrence of infimum gaps for strict-sense minimizers and abnormality in a higher-order Maximum Principle involving Lie brackets. While the abnormality-gap correspondence beyond first-order conditions has been already established for extended-sense --i.e. impulsive-- minimizers, a topological argument involving the former and the utilization of the above compatibility issues allow us to extend this correspondence to strict-sense minimizers.

math.OC↗

Higher-Order Normality and No-Gap Conditions in Impulsive Control with $L^1$-Control Topology

In optimal control, extending the class of admissible controls is a common strategy to guarantee the existence of optimal solutions. However, such extensions may introduce a gap between the infimum of the original problem and the minimum of the extended one, especially in the presence of endpoint constraints. Since Warga's seminal work, normality of first-order necessary conditions for extended minimizers has been recognized as a sufficient condition to avoid this phenomenon, though it is far from being necessary. In this paper, we consider impulsive extensions of control-affine systems with unbounded controls. We establish that a notion of \textit{higher-order normality}, based on iterated Lie brackets of the systems vector fields, suffices to prevent an infimum gap. The key novelty of this manuscript consists in showing that this holds under a local topology defined by the $L^1$-distance between controls, rather than the more common $L^\infty$-distance between trajectories. Among the reasons that motivate the interest in this issue, let us mention that a counterexample by R. B. Vinter shows that for a different extension -- based on convexification of the velocity set -- a local extended minimizer that is normal with respect to the $L^1$-norm of the controls may still exhibit a gap. Our method relies on set-separation techniques. Such an approach makes it possible to derive higher-order conditions and to exploit the corresponding notion of higher-order normality.

math.OC↗

Dynamic Programming Principle and Hamilton-Jacobi-Bellman Equation for Optimal Control Problems with Uncertainty

We study the properties of the value function associated with an optimal control problem with uncertainties, known as average or Riemann-Stieltjes problem. Uncertainties are assumed to belong to a compact metric probability space, and appear in the dynamics, in the terminal cost and in the initial condition, which yield an infinite-dimensional formulation. By stating the problem as an evolution equation in a Hilbert space, we show that the value function is the unique lower semi-continuous proximal solution of the Hamilton-Jacobi-Bellman (HJB) equation. Our approach relies on invariance properties and the dynamic programming principle.

math.OC↗

Online identification and control of PDEs via Reinforcement Learning methods

We focus on the control of unknown Partial Differential Equations (PDEs). The system dynamics is unknown, but we assume we are able to observe its evolution for a given control input, as typical in a Reinforcement Learning framework. We propose an algorithm based on the idea to control and identify on the fly the unknown system configuration. In this work, the control is based on the State-Dependent Riccati approach, whereas the identification of the model on Bayesian linear regression. At each iteration, based on the observed data, we obtain an estimate of the a-priori unknown parameter configuration of the PDE and then we compute the control of the correspondent model. We show by numerical evidence the convergence of the method for infinite horizon control problems.

math.OC↗

Convergence results for an averaged LQR problem with applications to reinforcement learning

In this paper, we will deal with a Linear Quadratic Optimal Control problem with unknown dynamics. As a modeling assumption, we will suppose that the knowledge that an agent has on the current system is represented by a probability distribution $π$ on the space of matrices. Furthermore, we will assume that such a probability measure is opportunely updated to take into account the increased experience that the agent obtains while exploring the environment, approximating with increasing accuracy the underlying dynamics. Under these assumptions, we will show that the optimal control obtained by solving the "average" Linear Quadratic Optimal Control problem with respect to a certain $π$ converges to the optimal control driven related to the Linear Quadratic Optimal Control problem governed by the actual, underlying dynamics. This approach is closely related to model-based Reinforcement Learning algorithms where prior and posterior probability distributions describing the knowledge on the uncertain system are recursively updated. In the last section, we will show a numerical test that confirms the theoretical results.

math.OC↗

Convergence of the Value Function in Optimal Control Problems with Unknown Dynamics

We deal with the convergence of the value function of an approximate control problem with uncertain dynamics to the value function of a nonlinear optimal control problem. The assumptions on the dynamics and the costs are rather general and we assume to represent uncertainty in the dynamics by a probability distribution. The proposed framework aims to describe and motivate some model-based Reinforcement Learning algorithms where the model is probabilistic. We also show some numerical experiments which confirm the theoretical results.

math.OC↗

Stability-Constrained Markov Decision Processes Using MPC

In this paper, we consider solving discounted Markov Decision Processes (MDPs) under the constraint that the resulting policy is stabilizing. In practice MDPs are solved based on some form of policy approximation. We will leverage recent results proposing to use Model Predictive Control (MPC) as a structured policy in the context of Reinforcement Learning to make it possible to introduce stability requirements directly inside the MPC-based policy. This will restrict the solution of the MDP to stabilizing policies by construction. The stability theory for MPC is most mature for the undiscounted MPC case. Hence, we will first show in this paper that stable discounted MDPs can be reformulated as undiscounted ones. This observation will entail that the MPC-based policy with stability requirements will produce the optimal policy for the discounted MDP if it is stable, and the best stabilizing policy otherwise.

cs.LG↗

Hamilton-Jacobi-Bellman Equation for Control Systems with Friction

This paper proposes a new framework to model control systems in which a dynamic friction occurs. The model consists in a controlled differential inclusion with a discontinuous right hand side, which still preserves existence and uniqueness of the solution for each given input function $u(t)$. Under general hypotheses, we are able to derive the Hamilton-Jacobi-Bellman equation for the related free time optimal control problem and to characterise the value function as the unique, locally Lipschitz continuous viscosity solution.

math.OC↗

Variational Problems for Tree Roots and Branches

This paper studies two classes of variational problems introduced in [7], related to the optimal shapes of tree roots and branches. Given a measure $μ$ describing the distribution of leaves, a sunlight functional $§(μ)$ computes the total amount of light captured by the leaves. For a measure $μ$ describing the distribution of root hair cells, a harvest functional $\H(μ)$ computes the total amount of water and nutrients gathered by the roots. In both cases, we seek a measure $μ$ that maximizes these functionals subject to a rami?ed transportation cost, for transporting nutrients from the roots to the trunk or from the trunk to the leaves. Compared with [7], here we do not impose any a priori bound on the total mass of the optimal measure $μ$, and more careful a priori estimates are thus required. In the unconstrained optimization problem for branches, we prove that an optimal measure exists, with bounded support and bounded total mass. In the unconstrained problem for tree roots, we prove that an optimal measure exists, with bounded support but possibly unbounded total mass. The last section of the paper analyzes how the size of the optimal tree depends on the parameters defining the various functionals.

math.OC↗

Necessary Conditions for Adverse Control Problems Expressed by Relaxed Derivatives

This paper provides a framework for deriving a new set of necessary conditions for adverse control problems among two players. The distinguish feature of such problems is that the first player has a priori knowledge on the second player strategy. A subclass of adverse control problems is the one of minimax control problems, which frequently arise in robust dynamic optimization. The conditions derived in this manuscript are expressed in terms of relaxed derivatives: the dual variables and the related functions are limits of computable sequences, obtained by considering a regularized version of the original problem and applying well known necessary condition. This topic was initially treated by J. Warga.

math.OC↗

Regularity of the Hamiltonian along Optimal Trajectories

This paper concerns state constrained optimal control problems, in which the dynamic constraint takes the form of a differential inclusion. If the differential inclusion does not depend on time, then the Hamiltonian, evaluated along the optimal state trajectory and the co-state trajectory, is independent of time. If the differential inclusion is Lipschitz continuous, then the Hamitonian, evaluated along the optimal state trajectory and the co-state trajectory, is Lipschitz continuous. These two well-known results are examples of the following principle: the Hamiltonian, evaluated along the optimal state trajectory and the co-state trajectory, inherits the regularity properties of the differential inclusion, regarding its time dependence. We show that this principle also applies to another kind of regularity: if the differential inclusion has bounded variation with respect to time, then the Hamiltonian, evaluated along the optimal state trajectory and the co-state trajectory, has bounded variation. Two applications of these newly found properties are demonstrated. One is to derive improved conditions which guarantee the nondegeneracy of necessary conditions of optimality, in the form of a Hamiltonian inclusion. The other application is to derive new, less restrictive, conditions under which minimizers in the calculus of variations have bounded slope. The analysis is based on a new, local, concept of differential inclusions that have bounded variation with respect to the time variable, in which conditions are imposed on the multifunction involved, only in a neighborhood of a given state trajectory.

math.OC↗

A geometrically based criterion to avoid infimum-gaps in Optimal Control

In optimal control theory the expression infimum gap means a strictly negative difference between the infimum value of a given minimum problem and the infimum value of a new problem obtained by the former by extending the original family V of controls to a larger family W. Now, for some classes of domain-extensions -- like convex relaxation or impulsive embedding of unbounded control problems -- the normality of an extended minimizer has been shown to be sufficient for the avoidance of an infimum gaps. A natural issue is then the search of a general hypothesis under which the criterium 'normality implies no gap' holds true. We prove that, far from being a peculiarity of those specific extensions and from requiring the convexity of the extended dynamics, this criterium is valid provided the original family V of controls is abundant in the extended family W. Abundance, which is stronger than the mere C^0-density of the original trajectories in the set of extended trajectories, is a dynamical-topological notion introduced by J. Warga, and is here utilized in a 'non-convex' version which, moreover, is adapted to differential manifolds. To get the main result, which is based on set separation arguments, we prove an open mapping result valid for Quasi-Differential-Quotient (QDQ) approximating cones, a notion of 'tangent cone' resulted as a peculiar specification of H. Sussmann's Approximate-Generalized-Differential-Quotients (AGDQ) approximating cone.

math.OC↗

A model for system uncertainty in reinforcement learning

This work provides a rigorous framework for studying continuous time control problems in uncertain environments. The framework considered models uncertainty in state dynamics as a measure on the space of functions. This measure is considered to change over time as agents learn their environment. This model can be seem as a variant of either Bayesian reinforcement learning or adaptive control. We study necessary conditions for locally optimal trajectories within this model, in particular deriving an appropriate dynamic programming principle and Hamilton-Jacobi equations. This model provides one possible framework for studying the tradeoff between exploration and exploitation in reinforcement learning.

math.OC↗

Growth Models for Tree Stems and Vines

The paper introduces a PDE model for the growth of a tree stem or a vine. The equations describe the elongation due to cell growth, and the response to gravity and to external obstacles. An additional term accounts for the tendency of a vine to curl around branches of other plants. When obstacles are present, the model takes the form of a differential inclusion with unilateral constraints. At each time t, a cone of admissible reactions is determined by the minimization of an elastic deformation energy. The main theorem shows that local solutions exist and can be prolonged globally in time, except when a specific "breakdown configuration" is reached. Approximate solutions are constructed by an operator-splitting technique. Some numerical simulations are provided at the end of the paper.

math.OC↗

A Stochastic Model of Optimal Debt Management and Bankruptcy

A problem of optimal debt management is modeled as a noncooperative game between a borrower and a pool of lenders, in infinite time horizon with exponential discount. The yearly income of the borrower is governed by a stochastic process. When the debt-to-income ratio $x(t)$ reaches a given size $x^*$, bankruptcy instantly occurs. The interest rate charged by the risk-neutral lenders is precisely determined in order to compensate for this possible loss of their investment. For a given bankruptcy threshold $x^*$, existence and properties of optimal feedback strategies for the borrower are studied, in a stochastic framework as well as in a limit deterministic setting. The paper also analyzes how the expected total cost to the borrower changes, depending on different values of $x^*$, changes, depending on different values of $x^*$?.

math.OC↗