SearcharxivSearch

arXiv subjects

Mehran Mesbahi

Publications and source records attributed to Mehran Mesbahi.

At least 19 recordsLinked to original sources

Unifying Variational View of Accelerated Primal-Dual Methods

We develop a unifying variational framework for accelerated primal-dual flows in affinely constrained convex optimization. In particular, it is shown that applying the Euler-Lagrange-Rayleigh (ELR) principle to coupled augmented Bregman Lagrangians systematically generates accelerated dynamics over general Bregman geometries and recovers a family of existing primal-dual mirror and Alternating Direction Method of Multipliers (ADMM) flows as special cases. We then establish $\mathcal{O}(e^{-b_t})$ convergence guarantees for these flows under mild assumptions. Lastly, we show that the proposed framework for algorithmic development extends from finite-dimensional optimization to constrained optimization over probability distributions.

math.OC

Hidden Star-Convexity in Policy Optimization for Gain-Scheduled LQR: Extended Version

We study policy optimization for gain-scheduled linear quadratic regulation, where one schedule of gains, interpolated through fixed weighting functions, is optimized against a family of plants. The resulting cost can develop spurious local minima, and existing convergence certificates are either local or severely conservative. We establish an exact identity: when the gradient of the cost is evaluated with the minimizer's closed-loop covariances, the scheduled cost is star-convex about the minimizer. The identity holds on the entire feasible set, for any parametrization of the schedule. Convergence is governed by a single dimensionless ratio. Wherever the ratio satisfies a threshold condition, gradient descent converges linearly to the optimum on entire sublevel regions at an explicit rate; at every spurious stationary point the condition necessarily fails. Experiments that maximize the ratio directly show the threshold to be an active boundary of the landscape. This extended version contains the complete proofs and additional numerical studies omitted from the letter for space.

math.OC

Dynamic Modeling and Parameter Estimation for Origami Structure Reconfiguration Process

The reconfiguration of origami during the folding and unfolding process is governed through a sequence of panel deformations and hinge orientations. To develop an effective model for representing the reconfiguration process, this paper introduces planar straight-line graphs and a novel consensus protocol for reaching the target origami configuration. The convergence and stability properties of the proposed consensus protocol are subsequently analyzed. Furthermore, to account for aggregate material and structural effects in the proposed consensus-based reconfiguration model, effective parameters embedded in the consensus protocol are identified from trajectory data using a fitting algorithm. Lastly, the effectiveness of the proposed modeling approach is shown using simulations of the two-panel structure and the Kresling origami pattern reconfiguration process.

math.DS

Receding-Horizon Policy Gradient for Polytopic Controller Synthesis

We propose the Polytopic Receding-Horizon Policy Gradient (P-RHPG) algorithm for synthesizing Parallel Distributed Compensation (PDC) controllers via Tensor Product (TP) model transformation. Standard LMI-based PDC synthesis grows increasingly conservative as model fidelity improves; P-RHPG instead solves a finite-horizon integrated cost via backward-stage decomposition. The key result is that each stage subproblem is a strongly convex quadratic in the vertex gains, a consequence of the linear independence of the HOSVD weighting functions, guaranteeing a unique global minimizer and linear convergence of gradient descent from any initialization. With zero terminal cost, the optimal cost increases monotonically to a finite limit and the gain sequence remains bounded; terminal costs satisfying a mild Lyapunov condition yield non-increasing convergence. Experiments on an aeroelastic wing benchmark confirm convergence to a unique infinite-horizon optimum across all tested terminal cost choices and near-optimal performance relative to the pointwise Riccati lower bound.

eess.SY

Robust Data-Driven Control for Nonlinear Systems Using their Digital Twins and Quadratic Funnels

This paper examines a robust data-driven approach for the safe deployment of systems with nonlinear dynamics using their imperfect digital twins. Our contribution involves proposing a method that fuses the digital twin's nominal trajectory with online, data-driven uncertainty quantification to synthesize robust tracking controllers. Specifically, we derive data-driven bounds to capture the deviations of the actual system from its prescribed nominal trajectory informed via its digital twin. Subsequently, the dataset is used in the synthesis of quadratic funnels -- robust positive invariant tubes around the nominal trajectory -- via linear matrix inequalities built on the time-series data. The resulting controller guarantees constraint satisfaction while adapting to the true system behavior through a segmented learning strategy, where each segment's controller is synthesized using uncertainty information from the previous segment. This work establishes a systematic framework for obtaining safety certificates in learning-based control of nonlinear systems with imperfect models.

eess.SY

Onboard Dual Quaternion Guidance for Rocket Landing

The dual quaternion guidance (DQG) algorithm was selected as the candidate 6-DoF powered-descent guidance algorithm for NASA's Safe and Precise Landing -- Integrated Capabilities Evolution (SPLICE) project. DQG is capable of handling state-triggered constraints that are of utmost importance in terms of enabling technologies such as terrain relative navigation. In this work, we develop a custom solver for DQG to enable onboard implementation for future rocket landing missions. We describe the design and implementation of a real-time-capable optimization framework, called sequential conic optimization (SeCO), that blends together sequential convex programming and first-order conic optimization to solve difficult nonconvex trajectory optimization problems, such as DQG, in real-time. A key feature of SeCO is that it leverages a first-order primal-dual conic optimization solver, based on the proportional-integral projected gradient method (PIPG). We describe the implementation of this solver, develop customizable first-order methods, and leverage convergence-accelerating strategies such as warm-starting and extrapolation, to solve the nonconvex DQG optimal control problem in real-time. Finally, in preparation for an upcoming closed-loop flight test campaign, we test our custom solver onboard the NASA SPLICE Descent and Landing Computer in a hardware-in-the-loop setting. We observe that our algorithm is significantly faster than previously reported solve-times using the flight-tested interior point method-based subproblem solver, BSOCP. Furthermore, our custom solver meets (and exceeds) NASA's autonomous precision rocket-landing guidance update-rate requirements for the first time, thus demonstrating the viability of SeCO for real-time, mission-critical applications onboard computationally-constrained flight hardware.

math.OC

Intrinsic Successive Convexification: Trajectory Optimization on Smooth Manifolds

A fundamental issue at the core of trajectory optimization on smooth manifolds is handling the implicit manifold constraint within the dynamics. The conventional approach is to enforce the dynamic model as a constraint. However, we show this approach leads to significantly redundant operations, as well as being heavily dependent on the state space representation. Specifically, we propose an intrinsic successive convexification methodology for optimal control on smooth manifolds. This so-called iSCvx is then applied to a representative example involving attitude trajectory optimization for a spacecraft subject to non-convex constraints.

math.OC

Estimation-Aware Trajectory Optimization with Set-Valued Measurement Uncertainties

In this paper, an optimization-based framework for generating estimation-aware trajectories is presented. In this setup, measurement (output) uncertainties are state-dependent and set-valued. Enveloping ellipsoids are employed to characterize state-dependent uncertainties with unknown distributions. The concept of regularity for set-valued output maps is then introduced, facilitating the formulation of the estimation-aware trajectory generation problem. Specifically, it is demonstrated that for output-regular maps, one can utilize a set-valued observability measure that is concave with respect to the finite horizon state trajectories. By maximizing this measure, estimation-aware trajectories can then be synthesized for a broad class of systems. Trajectory planning routines are also examined in this work, by which the observability measure is optimized for systems with locally linearized dynamics. To illustrate the effectiveness of the proposed approach, representative examples in the context of trajectory planning with vision-based estimation are presented. Moreover, the paper presents estimation-aware planning for an uncooperative Target-Rendezvous problem, where an Ego-satellite employs an onboard machine learning (ML)-based estimation module to realize the rendezvous trajectory.

math.OC

Multi Agent Reinforcement Learning for Sequential Satellite Assignment Problems

Assignment problems are a classic combinatorial optimization problem in which a group of agents must be assigned to a group of tasks such that maximum utility is achieved while satisfying assignment constraints. Given the utility of each agent completing each task, polynomial-time algorithms exist to solve a single assignment problem in its simplest form. However, in many modern-day applications such as satellite constellations, power grids, and mobile robot scheduling, assignment problems unfold over time, with the utility for a given assignment depending heavily on the state of the system. We apply multi-agent reinforcement learning to this problem, learning the value of assignments by bootstrapping from a known polynomial-time greedy solver and then learning from further experience. We then choose assignments using a distributed optimal assignment mechanism rather than by selecting them directly. We demonstrate that this algorithm is theoretically justified and avoids pitfalls experienced by other RL algorithms in this setting. Finally, we show that our algorithm significantly outperforms other methods in the literature, even while scaling to realistic scenarios with hundreds of agents and tasks.

cs.MA

Policy Optimization in Control: Geometry and Algorithmic Implications

This survey explores the geometric perspective on policy optimization within the realm of feedback control systems, emphasizing the intrinsic relationship between control design and optimization. By adopting a geometric viewpoint, we aim to provide a nuanced understanding of how various ``complete parameterization'' -- referring to the policy parameters together with its Riemannian geometry -- of control design problems, influence stability and performance of local search algorithms. The paper is structured to address key themes such as policy parameterization, the topology and geometry of stabilizing policies, and their implications for various (non-convex) dynamic performance measures. We focus on a few iconic control design problems, including the Linear Quadratic Regulator (LQR), Linear Quadratic Gaussian (LQG) control, and $\mathcal{H}_\infty$ control. In particular, we first discuss the topology and Riemannian geometry of stabilizing policies, distinguishing between their static and dynamic realizations. Expanding on this geometric perspective, we then explore structural properties of the aforementioned performance measures and their interplay with the geometry of stabilizing policies in presence of policy constraints; along the way, we address issues such as spurious stationary points, symmetries of dynamic feedback policies, and (non-)smoothness of the corresponding performance measures. We conclude the survey with algorithmic implications of policy optimization in feedback design.

math.OC

Six-Degree-of-Freedom Aircraft Landing Trajectory Planning with Runway Alignment

This paper presents a numerical optimization algorithm for generating approach and landing trajectories for a six-degree-of-freedom (6-DoF) aircraft. We improve on the existing research on aircraft landing trajectory generation by formulating the trajectory optimization problem with additional real-world operational constraints, including 6-DoF aircraft dynamics, runway alignment, constant wind field, and obstacle avoidance, to obtain a continuous-time nonconvex optimal control problem. Particularly, the runway alignment constraint enforces the trajectory of the aircraft to be aligned with the runway only during the final approach phase. This is a novel feature that is essential for preventing an approach that is either too steep or too shallow. The proposed method models the runway alignment constraint through a multi-phase trajectory planning scheme, imposing alignment conditions exclusively during the final approach phase. We compare this formulation with the existing state-triggered constraint formulation for runway alignment. To solve the formulated problem, we design a novel sequential convex programming algorithm called xPTR that extends the penalized trust-region (PTR) algorithm by incorporating an extrapolation step to expedite convergence. We validate the proposed method through extensive numerical simulations, including a Monte Carlo study, to evaluate the robustness of the algorithm to varying initial conditions.

math.OC

Output-feedback Synthesis Orbit Geometry: Quotient Manifolds and LQG Direct Policy Optimization

We consider direct policy optimization for the linear-quadratic Gaussian (LQG) setting. Over the past few years, it has been recognized that the landscape of dynamic output-feedback controllers of relevance to LQG has an intricate geometry, particularly pertaining to the existence of degenerate stationary points, that hinders gradient methods. In order to address these challenges, in this paper, we adopt a system-theoretic coordinate-invariant Riemannian metric for the space of dynamic output-feedback controllers and develop a Riemannian gradient descent for direct LQG policy optimization. We then proceed to prove that the orbit space of such controllers, modulo the coordinate transformation, admits a Riemannian quotient manifold structure. This geometric structure--that is of independent interest--provides an effective approach to derive direct policy optimization algorithms for LQG with a local linear rate convergence guarantee. Subsequently, we show that the proposed approach exhibits significantly faster and more robust numerical performance as compared with ordinary gradient descent.

math.OC

Data-Guided Regulator for Adaptive Nonlinear Control

This paper addresses the problem of designing a data-driven feedback controller for complex nonlinear dynamical systems in the presence of time-varying disturbances with unknown dynamics. Such disturbances are modeled as the "unknown" part of the system dynamics. The goal is to achieve finite-time regulation of system states through direct policy updates while also generating informative data that can subsequently be used for data-driven stabilization or system identification. First, we expand upon the notion of "regularizability" and characterize this system characteristic for a linear time-varying representation of the nonlinear system with locally-bounded higher-order terms. "Rapid-regularizability" then gauges the extent by which a system can be regulated in finite time, in contrast to its asymptotic behavior. We then propose the Data-Guided Regulation for Adaptive Nonlinear Control ( DG-RAN) algorithm, an online iterative synthesis procedure that utilizes discrete time-series data from a single trajectory for regulating system states and identifying disturbance dynamics. The effectiveness of our approach is demonstrated on a 6-DOF power descent guidance problem in the presence of adverse environmental disturbances.

eess.SY

An Active-Sensing Approach for Bearing-based Target Localization

Characterized by a cross-disciplinary nature, the bearing-based target localization task involves estimating the position of an entity of interest by a group of agents capable of collecting noisy bearing measurements. In this work, this problem is tackled by resting both on the weighted least square estimation approach and on the active-sensing control paradigm. Indeed, we propose an iterative algorithm that provides an estimate of the target position under the assumption of Gaussian noise distribution, which can be considered valid when more specific information is missing. Then, we present a seeker agents control law that aims at minimizing the localization uncertainty by optimizing the covariance matrix associated with the estimated target position. The validity of the designed bearing-based target localization solution is confirmed by the results of an extensive Monte Carlo simulation campaign.

eess.SY

Consensus on Lie groups for the Riemannian Center of Mass

In this paper, we develop a consensus algorithm for distributed computation of the Riemannian center of mass (RCM) on Lie Groups. The algorithm is built upon a distributed optimization reformulation that allows developing an intrinsic, distributed (without relying on a consensus subroutine), and a computationally efficient protocol for the RCM computation. The novel idea for developing this fast distributed algorithm is to utilize a Riemannian version of distributed gradient flow combined with a gradient tracking technique. We first guarantee that, under certain conditions, the limit point of our algorithm is the RCM point of interest. We then provide a proof of global convergence in the Euclidean setting, that can be viewed as a "geometric" dynamic consensus that converges to the average from arbitrary initial points. Finally, we proceed to showcase the superior convergence properties of the proposed approach as compared with other classes of consensus optimization-based algorithms for the RCM computation.

math.OC

Data-driven Optimal Filtering for Linear Systems with Unknown Noise Covariances

This paper examines learning the optimal filtering policy, known as the Kalman gain, for a linear system with unknown noise covariance matrices using noisy output data. The learning problem is formulated as a stochastic policy optimization problem, aiming to minimize the output prediction error. This formulation provides a direct bridge between data-driven optimal control and, its dual, optimal filtering. Our contributions are twofold. Firstly, we conduct a thorough convergence analysis of the stochastic gradient descent algorithm, adopted for the filtering problem, accounting for biased gradients and stability constraints. Secondly, we carefully leverage a combination of tools from linear system theory and high-dimensional statistics to derive bias-variance error bounds that scale logarithmically with problem dimension, and, in contrast to subspace methods, the length of output trajectories only affects the bias term.

eess.SY

Optimization-based Constrained Funnel Synthesis for Systems with Lipschitz Nonlinearities via Numerical Optimal Control

This paper presents a funnel synthesis algorithm for computing controlled invariant sets and feedback control gains around a given nominal trajectory for dynamical systems with locally Lipschitz nonlinearities and bounded disturbances. The resulting funnel synthesis problem involves a differential linear matrix inequality (DLMI) whose solution satisfies a Lyapunov condition that implies invariance and attractivity properties. Due to these properties, the proposed method can balance maximization of initial invariant funnel size, i.e., size of the funnel entry, and minimization of the size of the attractive funnel for attenuating the effect of disturbance. To solve the resulting funnel synthesis problem with the DLMI as constraints, we employ a numerical optimal control approach that uses a multiple shooting method to convert the problem into a finite dimensional semidefinite programming problem. This framework does not require piecewise linear system matrices and funnel parameters, which is typically assumed in recent related work. We illustrate the proposed funnel synthesis method with a numerical example.

math.OC

Duality-Based Stochastic Policy Optimization for Estimation with Unknown Noise Covariances

Duality of control and estimation allows mapping recent advances in data-guided control to the estimation setup. This paper formalizes and utilizes such a mapping to consider learning the optimal (steady-state) Kalman gain when process and measurement noise statistics are unknown. Specifically, building on the duality between synthesizing optimal control and estimation gains, the filter design problem is formalized as direct policy learning. In this direction, the duality is used to extend existing theoretical guarantees of direct policy updates for Linear Quadratic Regulator (LQR) to establish global convergence of the Gradient Descent (GD) algorithm for the estimation problem--while addressing subtle differences between the two synthesis problems. Subsequently, a Stochastic Gradient Descent (SGD) approach is adopted to learn the optimal Kalman gain without the knowledge of noise covariances. The results are illustrated via several numerical examples.

eess.SY