SearcharxivSearch

arXiv subjects

Mohammad Mahmoudi Filabadi

Publications and source records attributed to Mohammad Mahmoudi Filabadi.

3 recordsLinked to original sources

Unifying Entropy Regularization in Optimal Control: From and Back to Classical Objectives via Iterated Soft Policies and Path Integral Solutions

This paper develops a unified perspective on several optimal control formulations through the lens of Kullback-Leibler (KL) regularization. We propose a central problem that separates the KL penalties on policies and transitions with independent weights, thus generalizing the standard trajectory-level KL-regularization used in probabilistic optimal control. This umbrella formulation recovers various control problems: the classical Stochastic Optimal Control (SOC), Risk-Sensitive Stochastic Optimal Control (RSOC), and their policy-based KL-regularized counterparts, termed soft-policy SOC and RSOC, which yield tractable surrogates. Beyond being regularized variants, these soft-policy formulations majorize the original SOC and RSOC, thus, iterating their solutions recovers the original objectives. We further identify a synchronized case of soft-policy RSOC where the policy and transition KL weights coincide, yielding a linear Bellman operator, path-integral solution, and compositionality -- extending these computationally favourable properties to a broad class of control problems.

math.OC

Gaussian Process Dual MPC using Active Inference: An Autonomous Vehicle Usecase

Designing controllers under uncertainty requires balancing the need to explore system dynamics with the requirement to maintain reliable control performance. Dual control addresses this challenge by selecting actions that both regulate the system and actively gather informative data. This paper investigates the use of the Active Inference framework, grounded in the Free Energy Principle, for developing a dual model-predictive controller (MPC). To identify and quantify uncertainty, we introduce an online sparse semi-parametric Gaussian Process model that combines the flexibility of nonparametric with the efficiency of parametric learning for real-time updates. By applying the expected free energy functional to this adaptive probabilistic model, we derive an MPC objective that incorporates an information-theoretic term, which captures uncertainty arising from both the learned model and measurement noise. This formulation leads to a stochastic optimal control problem for dual controller design, which is solved using a novel dynamic-programming-based method. Simulation results on a vehicle use case demonstrate that the proposed algorithm enhances autonomous driving control performance across different settings and scenarios.

math.OC

Deterministic Trajectory Optimization through Probabilistic Optimal Control

In this article, we discuss two algorithms tailored to discrete-time deterministic finite-horizon nonlinear optimal control problems or so-called deterministic trajectory optimization problems. Both algorithms can be derived from an emerging theoretical paradigm that we refer to as probabilistic optimal control. The paradigm reformulates stochastic optimal control as an equivalent probabilistic inference problem and can be viewed as a generalisation of the former. The merit of this perspective is that it allows to address the problem using the Expectation-Maximization algorithm. It is shown that the application of this algorithm results in a fixed point iteration of probabilistic policies that converge to the deterministic optimal policy. Two strategies for policy evaluation are discussed, using state-of-the-art uncertainty quantification methods resulting into two distinct algorithms. The algorithms are structurally closest related to the differential dynamic programming algorithm and related methods that use sigma-point methods to avoid direct gradient evaluations. The main advantage of the algorithms is an improved balance between exploration and exploitation over the iterations, leading to improved numerical stability and accelerated convergence. These properties are demonstrated on different nonlinear systems.

math.OC