SearcharxivSearch

arXiv subjects

Mario Zanon

Publications and source records attributed to Mario Zanon.

At least 19 recordsLinked to original sources

Characterization and Computation of Feedback Nash Equilibria in Scalar Discounted N-Player Linear Quadratic Games

This paper studies feedback Nash equilibria (FNE) in scalar discounted linear quadratic (LQ) games with $N$ players. By explicitly incorporating the discount factor, we show that finite-cost equilibria may fail to stabilize the original system, motivating a distinction between FNE and stable FNE together with a sufficient stability condition. Based on a parametric characterization of the policies, we propose numerical methods for computing all equilibria. Particular attention is devoted to the symmetric game, where a closed-form expression of the symmetric FNE and conditions for the existence of up to $M\leq2^N-2$ equilibria are derived. Numerical experiments illustrate how equilibrium multiplicity depends on the game configuration and highlight the emergence of finite-cost non-stabilizing equilibria.

eess.SY

On Piecewise Quadratic Terminal Costs for MPC

This paper presents a novel approach to synthesize stabilizing termi- nal ingredients for linear model predictive control (MPC) schemes, with the aim of increasing the region of attraction while reducing suboptimal- ity with respect to the solution of the infinite-horizon optimal control problem. It is based on the construction of a novel terminal region using methods from the field of configuration-constrained polytopic computing, along with a terminal cost that is exactly equal to the infinite-horizon linear-quadratic regulator cost in a nontrivial neighborhood of the steady- state. The practical performance of the controller is illustrated through various case studies, and comparisons with state-of-the-art approaches are presented.

eess.SY

Active Learning MPC Objective Functions from Preferences

Designing the objective function in Model Predictive Control (MPC) is challenging when performance assessment criteria are available only from human judgment. We adopt a preference-based learning (PbL) approach to learn the MPC objective function from preferences over trajectory pairs. However, the real-world application of PbL is often restricted by the significant cost or limited availability of human preference queries. To address this, Active Learning (AL) strategies seek to improve sampling efficiency, reducing the labeling effort required to obtain a well-performing classifier. We present two AL strategies for learning the MPC objective function from human preferences over pairwise system trajectories: a pool-based strategy that selects trajectory pairs that are both uncertain under the current surrogate and diverse relative to previously labeled comparisons, and a query-synthesis strategy that incorporates new trajectories using the current surrogate-driven MPC. Numerical results show that the proposed strategies yield closed-loop behaviors that align more with the expressed preference using fewer number of queries compared to a random sampling approach.

eess.SY

Fast Gauss-Newton for Multiclass Cross-Entropy

In multiclass softmax cross-entropy, the full generalized Gauss-Newton (GGN) curvature couples all output logits through the softmax covariance, making curvature-vector products harder to scale as the number of classes grows. We show that the standard multiclass GGN can be decomposed exactly into a true-vs-rest term and a positive semidefinite within-competitor covariance term. Fast Gauss-Newton (FGN) retains the first term and drops the second, yielding a positive semidefinite under-approximation of the multiclass GGN that is exact for binary classification. The derivation uses an exact true-vs-rest scalar-margin representation of softmax cross-entropy: the loss and gradient are unchanged, and the approximation enters only at the curvature level. Exploiting the FGN curvature structure, the damped update can be written as an equivalent whitened row-space system with one row per mini-batch example. We solve this system matrix-free by conjugate gradient using Jacobian-vector and vector-Jacobian products of the scalar margin map. Targeted mechanism experiments and an evaluation on a fixed-feature multiclass head support the predictions from the decomposition: FGN stays closest to the full softmax GGN when competitor mass is concentrated or damping is large, and deviates as the dropped within-competitor covariance grows.

cs.LG

Second-Order, First-Class: A Composable Stack for Curvature-Aware Training

Second-order methods promise improved stability and faster convergence, yet they remain underused due to implementation overhead, tuning brittleness, and the lack of composable APIs. We introduce Somax, a composable Optax-native stack that treats curvature-aware training as a single JIT-compiled step governed by a static plan. Somax exposes first-class modules -- curvature operators, estimators, linear solvers, preconditioners, and damping policies -- behind a single step interface and composes with Optax by applying standard gradient transformations (e.g., momentum, weight decay, schedules) to the computed direction. This design makes typically hidden choices explicit and swappable. Somax separates planning from execution: it derives a static plan (including cadences) from module requirements, then runs the step through a specialized execution path that reuses intermediate results across modules. We report system-oriented ablations showing that (i) composition choices materially affect scaling behavior and time-to-accuracy, and (ii) planning reduces per-step overhead relative to unplanned composition with redundant recomputation.

cs.LG

On Two-Player Scalar Discrete-Time Linear Quadratic Games

For the characterization of Feedback Nash Equilibria (FNE) in linear quadratic games, this paper provides a detailed analysis of the discrete-time discounted coupled best-response equations for the scalar two-player setting, together with a set of analytical tools for the classification of local saddle property for the iterative best-response method. Through analytical and numerical results we show the importance of classification, revealing an anti-coordination scheme in the case of multiple solutions. Particular attention is given to the symmetric case, where identical cost function parameters allow closed-form expressions and explicit necessary and sufficient conditions for the existence and multiplicity of FNE. We also present numerical results that illustrate the theoretical findings and offer foundational insights for the design and validation of iterative NE-seeking methods.

math.OC

Rethinking Strict Dissipativity for Economic MPC

Stability of economic model predictive control can be proven under the assumption that a strict dissipativity condition holds. This assumption has a clear interpretation in terms of the so-called rotated stage cost, which must have its minimum at the optimal steady state. However, contrary to dissipativity, for strict dissipativity the storage function cannot be immediately related to the value function of an optimal control problem formulated with the economic stage cost. We propose the novel concept of two-storage strict dissipativity, which requires two storage functions to satisfy dissipativity and be separated by a positive definite function. This new condition can be immediately related to optimal control by means of value functions and might be easier to verify than strict dissipativity. Furthermore, we prove that two-storage strict dissipativity is sufficient and necessary for asymptotic stability, it is related to strict dissipativity, and also to alternative approaches relying on the so-called cost-to-travel. Finally, we discuss commonly used and new terminal cost designs that guarantee asymptotic stability in the finite-horizon case.

math.OC

Learning the MPC objective function from human preferences

In Model Predictive Control (MPC), the objective function plays a central role in determining the closed-loop behavior of the system, and must therefore be designed to achieve the desired closed-loop performance. However, in real-world scenarios, its design is often challenging, as it requires balancing complex trade-offs and accurately capturing a performance criterion that may not be easily quantifiable in terms of an objective function. This paper explores preference-based learning as a data-driven approach to constructing an objective function from human preferences over trajectory pairs. We formulate the learning problem as a machine learning classification task to learn a surrogate model that estimates the likelihood of a trajectory being preferred over another. The approach provides a surrogate model that can directly be used as an MPC objective function. Numerical results show that we can learn objective functions that provide closed-loop trajectories that align with the expressed human preferences.

eess.SY

Economic Linear Quadratic MPC With Non-Unique Optimal Solutions

Asymptotic stability in economic receding horizon control can be obtained under a strict dissipativity assumption, related to positive-definiteness of a so-called rotated cost, and through the use of suitable terminal cost and constraints. In the linear-quadratic case a common assumption is that the rotated cost is positive definite. The positive semi-definite case has received surprisingly little attention, and the connection to the standard dissipativity assumption has not been investigated. In this paper, we fill this gap by connecting existing results in economic model predictive control with the stability results for the semi-definite case, the properties of the constrained generalized discrete algebraic Riccati equation, and of two optimal control problems. Moreover, we extend recent results relating exponential stability to the choice of terminal cost in the absence of terminal constraints.

eess.SY

Stabilization of Strictly Pre-Dissipative Receding Horizon Linear Quadratic Control by Terminal Costs

Asymptotic stability in receding horizon control is obtained under a strict pre-dissipativity assumption, in the presence of suitable state constraints. In this paper we analyze how terminal constraints can be replaced by suitable terminal costs. We restrict to the linear-quadratic setting as that allows us to obtain stronger results, while we analyze the full nonlinear case in a separate contribution.

math.OC

Optimality Conditions for Model Predictive Control: Rethinking Predictive Model Design

Optimality is a critical aspect of Model Predictive Control (MPC), especially in economic MPC. However, achieving optimality in MPC presents significant challenges, and may even be impossible, due to inherent inaccuracies in the predictive models. Predictive models often fail to accurately capture the true system dynamics, such as in the presence of stochasticity, leading to suboptimal MPC policies. In this paper, we establish the necessary and sufficient conditions on the underlying prediction model for an MPC scheme to achieve closed-loop optimality. Interestingly, these conditions are counterintuitive to the traditional approach of building predictive models that best fit the data. These conditions present a mathematical foundation for constructing models that are directly linked to the performance of the resulting MPC scheme.

math.OC

Incremental Gauss-Newton Descent for Machine Learning

Stochastic gradient updates are widely used for their efficiency and scalability, but their effective step sizes can depend strongly on feature scaling and local model sensitivity. Gauss-Newton methods address such scale effects through curvature information, but in their standard mini-batch form they require matrix-vector products, linear solves, or structured approximations. This paper studies the special case of scalar-output losses evaluated one sample at a time. In this setting, the generalized Gauss-Newton matrix has rank at most one, and its only possible nonzero curvature direction is aligned with the stochastic gradient. As a result, the damped Gauss-Newton direction reduces to a closed-form scalar normalization of the sample gradient. The resulting update, Incremental Gauss-Newton Descent (IGND), requires no curvature matrix storage, factorization, or iterative linear solve. We derive the update, characterize its behavior, and relate it to normalized gradient descent, adaptive first-order methods, stochastic Polyak step sizes, and mini-batch Gauss-Newton updates. Under explicit smoothness, alignment, and stochastic approximation assumptions, we prove a stationarity result for the IGND update. Experiments on supervised learning, a controlled test of scale robustness, and a linear-quadratic control case study show that IGND improves robustness to sensitivity scaling and can be competitive with, or complementary to, common stochastic optimizers while retaining a simple incremental update.

cs.LG

Exact Gauss-Newton Optimization for Training Deep Neural Networks

We present Exact Gauss-Newton (EGN), a stochastic second-order optimization algorithm that combines the generalized Gauss-Newton (GN) Hessian approximation with low-rank linear algebra to compute the descent direction. Leveraging the Duncan-Guttman matrix identity, the parameter update is obtained by factorizing a matrix which has the size of the mini-batch. This is particularly advantageous for large-scale machine learning problems where the dimension of the neural network parameter vector is several orders of magnitude larger than the batch size. Additionally, we show how improvements such as line search, adaptive regularization, and momentum can be seamlessly added to EGN to further accelerate the algorithm. Moreover, under mild assumptions, we prove that our algorithm converges in expectation to a stationary point of the objective. Finally, our numerical experiments demonstrate that EGN consistently exceeds, or at most matches the generalization performance of well-tuned SGD, Adam, GAF, SQN, and SGN optimizers across various supervised and reinforcement learning tasks.

cs.LG

Optimization-based Heuristic for Vehicle Dynamic Coordination in Mixed Traffic Intersections

In this paper, we address a coordination problem for connected and autonomous vehicles (CAVs) in mixed traffic settings with human-driven vehicles (HDVs). The main objective is to have a safe and optimal crossing order for vehicles approaching unsignalized intersections. This problem results in a mixed-integer quadratic programming (MIQP) formulation which is unsuitable for real-time applications. Therefore, we propose a computationally tractable optimization-based heuristic that monitors platoons of CAVs and HDVs to evaluate whether alternative crossing orders can perform better. It first checks the future constraint violation that consistently occurs between pairs of platoons to determine a potential swap. Next, the costs of quadratic programming (QP) formulations associated with the current and alternative orders are compared in a depth-first branching fashion. In simulations, we show that the heuristic can be a hundred times faster than the original and simplified MIQPs and yields solutions that are close to optimal and have better order consistency.

eess.SY

Learning disturbance models for offset-free reference tracking

This work presents a nonlinear control framework that guarantees asymptotic offset-free tracking of generic reference trajectories by learning a nonlinear disturbance model, which compensates for input disturbances and model-plant mismatch. Our approach generalizes the well-established method of using an observer to estimate a constant disturbance to allow tracking constant setpoints with zero steady-state error. In this paper, the disturbance model is generalized to a nonlinear static function of the plant's state and command input, learned online, so as to perfectly track time-varying reference trajectories under certain assumptions on the model and provided that future reference samples are available. We compare our approach with the classical constant disturbance model in numerical simulations, showing its superiority.

eess.SY

Computation of safe disturbance sets using implicit RPI sets

Given a stable linear time-invariant (LTI) system subject to output constraints, we present a method to compute a set of disturbances such that the reachable set of outputs matches as closely as possible the output constraint set, while being included in it. This problem finds application in several control design problems, such as the development of hierarchical control loops, decentralized control, supervisory control, robustness-verification, etc. We first characterize the set of disturbance sets satisfying the output constraint inclusion using corresponding minimal robust positive invariant (mRPI) sets, following which we formulate an optimization problem that minimizes the distance between the reachable output set and the output constraint set. We tackle the optimization problem using an implicit RPI set approach that provides a priori approximation error guarantees, and adopt a novel disturbance set parameterization that permits the encoding of the set of feasible disturbance sets as a polyhedron. Through extensive numerical examples, we demonstrate that the proposed approach computes disturbance sets with reduced conservativeness improved computational efficiency than state-of-the-art methods.

eess.SY

Experimental Validation of Safe MPC for Autonomous Driving in Uncertain Environments

The full deployment of autonomous driving systems on a worldwide scale requires that the self-driving vehicle be operated in a provably safe manner, i.e., the vehicle must be able to avoid collisions in any possible traffic situation. In this paper, we propose a framework based on Model Predictive Control (MPC) that endows the self-driving vehicle with the necessary safety guarantees. In particular, our framework ensures constraint satisfaction at all times, while tracking the reference trajectory as close as obstacles allow, resulting in a safe and comfortable driving behavior. To discuss the performance and real-time capability of our framework, we provide first an illustrative simulation example, and then we demonstrate the effectiveness of our framework in experiments with a real test vehicle.

cs.RO