SearcharxivSearch

arXiv subjects

Adwait Datar

Publications and source records attributed to Adwait Datar.

12 recordsLinked to original sources

Well-Posed KL-Regularized Control via Wasserstein and Kalman-Wasserstein KL Divergences

Kullback-Leibler (KL) divergence regularization is widely used in reinforcement learning, but it becomes infinite under support mismatch and can degenerate in low-noise regimes. Using a unified information-geometric framework, we introduce KL analogs by replacing the Fisher-Rao geometry in the dynamical formulation of the KL with transport-based geometries, and derive closed-form expressions for common distribution families. Between elliptic distributions, these divergences remain finite for degenerating equal covariances and yield a geometric interpretation of regularization heuristics used in Kalman ensemble methods. We demonstrate the utility of these divergences in KL-regularized optimal control. In the fully tractable setting of linear time-invariant systems with Gaussian process noise, the classical KL reduces to a quadratic control penalty that becomes singular as process noise vanishes. Our variants remove this singularity and yield well-posed problems. In both the double integrator and cart-pole examples, the resulting controls preserve nontrivial feedback and achieve better closed-loop performance.

math.OC

Convergence Properties of Natural Gradient Descent for Minimizing KL Divergence

The Kullback-Leibler (KL) divergence plays a central role in probabilistic machine learning, where it commonly serves as the canonical loss function. Optimization in such settings is often performed over the probability simplex, where the choice of parameterization significantly impacts convergence. In this work, we study the problem of minimizing the KL divergence and analyze the behavior of gradient-based optimization algorithms under two dual coordinate systems within the framework of information geometry$-$ the exponential family ($\theta$ coordinates) and the mixture family ($\eta$ coordinates). We compare Euclidean gradient descent (GD) in these coordinates with the coordinate-invariant natural gradient descent (NGD), where the natural gradient is a Riemannian gradient that incorporates the intrinsic geometry of the underlying statistical model. In continuous time, we prove that the convergence rates of GD in the $\theta$ and $\eta$ coordinates provide lower and upper bounds, respectively, on the convergence rate of NGD. Moreover, under affine reparameterizations of the dual coordinates, the convergence rates of GD in $\eta$ and $\theta$ coordinates can be scaled to $2c$ and $\frac{2}{c}$, respectively, for any $c>0$, while NGD maintains a fixed convergence rate of $2$, remaining invariant to such transformations and sandwiched between them. Although this suggests that NGD may not exhibit uniformly superior convergence in continuous time, we demonstrate that its advantages become pronounced in discrete time, where it achieves faster convergence and greater robustness to noise, outperforming GD. Our analysis hinges on bounding the spectrum and condition number of the Hessian of the KL divergence at the optimum, which coincides with the Fisher information matrix.

cs.LG

Information Geometry for Wasserstein KL Divergence of Gaussian Measures on $\mathbb{R}^n$

We study the Wasserstein Kullback--Leibler divergence (WKL divergence) on the manifold of nondegenerate Gaussian measures over $\mathbb R^n$. In the canonical-divergence construction, the Fisher--Rao metric recovers forward KL along intrinsic mixture geodesics and reverse KL along geodesics of the conjugate exponential connection. Replacing the Fisher--Rao metric by the Otto metric and following the latter route produces the $e_1$-connection underlying WKL divergence. We establish its geodesic completeness, classify its forward limits, prove that every ordered pair is joined by a unique $e_1$-connector generated by a quadratic potential, and derive an explicit WKL divergence formula with separate mean and covariance contributions. WKL divergence is nonnegative and separating. At equal covariances WKL divergence equals one half of the squared Euclidean mean distance. Finally, WKL divergence extends finitely and continuously to singular targets from a nondegenerate source, but diverges when the target remains nondegenerate and the source covariance becomes singular. Path-dependent joint limits at Dirac pairs preclude a continuous extension to the full positive-semidefinite covariance product, although a lower-semicontinuous extended-real extension exists.

math.ST

Robust Stability for Multiagent Systems with Spatio-Temporally Correlated Packet Loss

A problem with considering correlations in the analysis of multiagent system with stochastic packet loss is that they induce dependencies between agents that are otherwise decoupled, preventing the application of decomposition methods required for efficient evaluation. To circumvent that issue, this paper is proposing an approach based on analysing sets of networks with independent communication links, only considering the correlations in an implicit fashion. Combining ideas from the robust stabilization of Markov jump linear systems with recently proposed techniques for analysing packet loss in multiagent systems, we obtain a linear matrix inequality based stability condition which is independent of the number of agents. The main result is that the set of stabilized probability distributions has non-empty interior such that small correlations cannot lead to instability, even though only distributions of independent links were analysed. Moreover, two examples are provided to demonstrate the applicability of the results to practically relevant scenarios.

math.OC

Systematic construction of continuous-time neural networks for linear dynamical systems

Discovering a suitable neural network architecture for modeling complex dynamical systems poses a formidable challenge, often involving extensive trial and error and navigation through a high-dimensional hyper-parameter space. In this paper, we discuss a systematic approach to constructing neural architectures for modeling a subclass of dynamical systems, namely, Linear Time-Invariant (LTI) systems. We use a variant of continuous-time neural networks in which the output of each neuron evolves continuously as a solution of a first-order or second-order Ordinary Differential Equation (ODE). Instead of deriving the network architecture and parameters from data, we propose a gradient-free algorithm to compute sparse architecture and network parameters directly from the given LTI system, leveraging its properties. We bring forth a novel neural architecture paradigm featuring horizontal hidden layers and provide insights into why employing conventional neural architectures with vertical hidden layers may not be favorable. We also provide an upper bound on the numerical errors of our neural networks. Finally, we demonstrate the high accuracy of our constructed networks on three numerical examples.

cs.LG

Transformation-Free Fixed-Structure Model Reduction for LPV Systems

In this paper, we propose a model reduction technique for linear parameter varying (LPV) systems based on available tools for fixed-structure controller synthesis. We start by transforming a model reduction problem into an equivalent controller synthesis problem by defining an appropriate generalized plant. The controller synthesis problem is then solved by using gradient-based tools available in the literature. Owing to the flexibility of the gradient-based synthesis tools, we are able to impose a desired structure on the obtained reduced model. Additionally, we obtain a bound on the approximation error as a direct output of the optimization problem. The proposed methods are applied on a benchmark mechanical system of interconnected masses, springs and dampers. To evaluate the effect of the proposed model-reduction approach on controller design, LPV controllers designed using the reduced models (with and without an imposed structure) are compared in closed-loop with the original model.

eess.SY

On the Natural Gradient of the Evidence Lower Bound

This article studies the Fisher-Rao gradient, also referred to as the natural gradient, of the evidence lower bound (ELBO) which plays a central role in generative machine learning. It reveals that the gap between the evidence and its lower bound, the ELBO, has essentially a vanishing natural gradient within unconstrained optimization. As a result, maximization of the ELBO is equivalent to minimization of the Kullback-Leibler divergence from a target distribution, the primary objective function of learning. Building on this insight, we derive a condition under which this equivalence persists even when optimization is constrained to a model. This condition yields a geometric characterization, which we formalize through the notion of a cylindrical model.

cs.LG

Gradient-based Cooperative Control of quasi-Linear Parameter Varying Vehicles with Noisy Gradients

This paper extends recent results on the exponential performance analysis of gradient based cooperative control dynamics using the framework of exponential integral quadratic constraints ($\alpha-$IQCs). A cooperative source-seeking problem is considered as a specific example where one or more vehicles are embedded in a strongly convex scalar field and are required to converge to a formation located at the minimum of a field. A subset of the agents are assumed to have the knowledge of the gradient of the field evaluated at their respective locations and the interaction graph is assumed to be uncertain. As a first contribution, we extend earlier results on linear time invariant (LTI) systems to non-linear systems by using quasi-linear parameter varying (qLPV) representations. Secondly, we remove the assumption on perfect gradient measurements and consider multiplicative noise in the analysis. Performance-robustness trade off curves are presented to illustrate the use of presented methods for tuning controller gains. The results are demonstrated on a non-linear second order vehicle model with a velocity-dependent non-linear damping and a local gain-scheduled tracking controller.

math.OC

A Decomposition Approach to Multi-Agent Systems with Bernoulli Packet Loss

In this paper, we extend the decomposable systems framework to multi-agent systems with Bernoulli distributed packet loss with uniform probability. The proposed sufficient analysis conditions for mean-square stability and $H_2$-performance -- which are expressed in the form of linear matrix inequalities -- scale linearly with increased network size and thus allow to analyse even very large-scale multi-agent systems. A numerical example demonstrates the potential of the approach by application to a first-order consensus problem.

math.OC

Robust Performance Analysis of Cooperative Control Dynamics via Integral Quadratic Constraints

We study cooperative control dynamics with gradient based forcing terms. As a specific example, we focus on source-seeking dynamics with vehicles embedded in an unknown scalar field with a subset of agents having gradient information. As interaction mechanisms, formation control dynamics and flocking dynamics are considered. We leverage the framework of $\alpha$-integral quadratic constraints to obtain convergence rate estimates whenever exponential stability can be achieved. The communication graph and the interaction potential are assumed to be time-invariant and uncertain. Sufficient conditions take the form of linear matrix inequalities independent of the size of network. A derivation (purely in time-domain) of the so-called \textit{hard} Zames-Falb $\alpha$-IQCs involving general non-causal higher order multipliers is given along with a suitably adapted parameterization of the multipliers to the $\alpha$-IQC setting. The time-domain arguments facilitate a straightforward extension to linear parameter varying systems. Numerical examples illustrate the application of the theoretical results.

math.OC

Finite State Markov Modeling of C-V2X Erasure Links For Performance and Stability Analysis of Platooning Applications

Cooperative driving systems, such as platooning, rely on communication and information exchange to create situational awareness for each agent. Design and performance of control components are therefore tightly coupled with communication component performance. The information flow between vehicles can significantly affect the dynamics of a platoon. Therefore, both the performance and the stability of a platoon depend not only on the vehicle's controller but also on the information flow Topology (IFT). The IFT can cause limitations for certain platoon properties, i.e., stability and scalability. Cellular Vehicle-To-Everything (C-V2X) has emerged as one of the main communication technologies to support connected and automated vehicle applications. As a result of packet loss, wireless channels create random link interruption and changes in network topologies. In this paper, we model the communication links between vehicles with a first-order Markov model to capture the prevalent time correlations for each link. These models enable performance evaluation through better approximation of communication links during system design stages. Our approach is to use data from experiments to model the Inter-Packet Gap (IPG) using Markov chains and derive transition probability matrices for consecutive IPG states. Training data is collected from high fidelity simulations using models derived based on empirical data for a variety of different vehicle densities and communication rates. Utilizing the IPG models, we analyze the mean-square stability of a platoon of vehicles with the standard consensus protocol tuned for ideal communication and compare the degradation in performance for different scenarios.

cs.RO

Robust Performance Analysis of Source-Seeking Dynamics with Integral Quadratic Constraints

We analyze the performance of source-seeking dynamics involving either a single vehicle or multiple flocking-vehicles embedded in an underlying strongly convex scalar field with gradient based forcing terms. For multiple vehicles under flocking dynamics embedded in quadratic fields, we show that the dynamics of the center of mass are equivalent to the dynamics of a single agent. We leverage the recently developed framework of $\alpha$-integral quadratic constraints (IQCs) to obtain convergence rate estimates. We first present a derivation of \textit{hard} Zames-Falb (ZF) $\alpha$-IQCs involving general non-causal multipliers based on purely time-domain arguments and show that a parameterization of the ZF multiplier, suggested in the literature for the standard version of the ZF IQCs, can be adapted to the $\alpha$-IQCs setting to obtain quasi-convex programs for estimating convergence rates. Owing to the time-domain arguments, we can seamlessly extend these results to linear parameter varying (LPV) vehicles possibly opening the doors to non-linear vehicle models with quasi-LPV representations. We illustrate the theoretical results on a linear time invariant (LTI) model of a quadrotor, a non-minimum phase LTI plant and two LPV examples which show a clear benefit of using general non-causal dynamic multipliers to drastically reduce conservatism.

math.OC