SearcharxivSearch

arXiv subjects

Somnath Pradhan

Publications and source records attributed to Somnath Pradhan.

At least 19 recordsLinked to original sources

On the Approximation of Optimal Control in Regime-Switching Diffusions

We study approximation and structural simplification of optimal control policies for controlled regime-switching diffusion processes for discounted, ergodic, finite-horizon, and exit-time criteria. We first establish continuity of the cost functionals over classes of Markov and stationary Markov policies by exploiting elliptic and parabolic regularity of the corresponding Hamilton--Jacobi--Bellman and Poisson equations. Using density results of policies with finite-action, piecewise-constant, and Lipschitz continuous, we show that each control problem admits $\varepsilon$-optimal policies within these structured subclasses. We then construct an Euler--Maruyama approximation of the controlled regime-switching diffusion under piecewise-constant controls. We prove strong convergence of the controlled state process and establish convergence of the associated finite-horizon value functions with rate $O(h^{\gamma/2})$. Building on this discretization, we develop a finite-state approximation of the induced discrete-time Markov chain via state-space quantization. We show that the value functions of the finite models converge uniformly on compact sets to the value function of the original problem, and that optimal policies of the approximating models are asymptotically optimal. These results provide a systematic framework for approximating regime-switching diffusion control problems and justify the use of structured policies and finite-state models for numerical implementation.

math.OC

Near Optimality of Discrete-Time Approximations for Controlled McKean-Vlasov and Large Interacting Particle Diffusions

We study stochastic optimal control problems for (possibly degenerate) McKean-Vlasov controlled diffusions and obtain discrete-time as well as finite interacting particle approximations. (i) Via continuity of the expected cost in control policy by endowing the space of relaxed policies with a compact weak topology, we prove near-optimality of piecewise-constant policies which leads to a discrete-time model. We show that the discrete-time value functions (for finite-horizon and discounted infinite-horizon) converge to their continuous-time counterparts as the timestep converges to zero. In particular, we establish that optimal policies for the discrete-time model exists and they are near-optimal for the original continuous-time problem. (ii) We then extend these approximation and near-optimality results to $N$-particle interacting systems under centralized or decentralized mean-field sharing information structure, proving that the discrete-time McKean-Vlasov policy is asymptotically optimal as $N\to \infty$ and the time step goes to zero. Using discrete-time approximation as an intermediate step leads to complementary conditions compared to those in the literature. (iii) We thus develop a unified approximation framework for McKean-Vlasov optimal control problems via discrete-time McKean-Vlasov control problems (and associated numerical methods such as finite model approximations), and we also show near optimality of such approximate policy solutions for the $N$-agent interacting models under centralized and decentralized control.

math.OC

Reinforcement Learning for Discounted and Ergodic Control of Diffusion Processes

This paper develops a quantized Q-learning algorithm for the optimal control of controlled diffusion processes on $\mathbb{R}^d$ under both discounted and ergodic (average) cost criteria. We first establish near-optimality of finite-state MDP approximations to discrete-time discretizations of the diffusion, then introduce a quantized Q-learning scheme and prove its almost-sure convergence to near-optimal policies for the finite MDP. These policies, when interpolated to continuous time, are shown to be near-optimal for the original diffusion model under discounted costs and -- via a vanishing-discount argument -- also under ergodic costs for sufficiently small discount factors. The analysis applies under mild conditions (Lipschitz dynamics, non-degeneracy, bounded continuous costs, and Lyapunov stability for ergodic case) without requiring prior knowledge of the system dynamics or restrictions on control policies (beyond admissibility). Our results complement recent work on continuous-time reinforcement learning for diffusions by providing explicit near-optimality rates and extending rigorous guarantees both for discounted cost and ergodic cost criteria for diffusions with unbounded state space.

math.OC

Robustness of optimal control for controlled regime-switching diffusions with incorrect models

This paper investigates the robustness of stochastic optimal control for controlled regime switching diffusions. We consider systems driven by both continuous fluctuations and discrete regime changes, allowing for model misspecification in both the diffusion and switching components. Within a unified framework, we study four classical cost formulations finite horizon, infinite-horizon discounted and ergodic costs, and the exit time cost, and establish continuity of value functions and robustness of optimal controls. Specifically, we show that as a sequence of approximating regime switching models converges to the true model, the associated value functions and optimal policies converge as well, ensuring vanishing performance loss. The analysis relies on the regularity of the solution to the associated weakly coupled HJB systems, and their stochastic representation. The results extend the robustness framework developed for diffusion processes to a significantly broader class of hybrid systems with interacting continuous and discrete dynamics.

math.OC

Controlled Diffusions under Full, Partial and Decentralized Information: Existence of Optimal Policies and Discrete-Time Approximations

We present existence and discrete-time approximation results on optimal control policies for continuous-time stochastic control problems under a variety of information structures. These include fully observed models, partially observed models and multi-agent models with decentralized information structures. While there exist comprehensive existence and approximations results for the fully observed setup in the literature, few prior research exists on discrete-time approximation results for partially observed models. For decentralized models, even existence results have not received much attention except for specialized models and approximation has been an open problem. Our existence and approximations results lead to the applicability of well-established partially observed Markov decision processes and the relatively more mature theory of discrete-time decentralized stochastic control to be applicable for computing near optimal solutions for continuous-time stochastic control.

math.OC

Discrete-Time Approximations of Controlled Diffusions with Infinite Horizon Discounted and Average Cost

We present discrete-time approximation of optimal control policies for infinite horizon discounted/ergodic control problems for controlled diffusions in $\Rd$\,. In particular, our objective is to show near optimality of optimal policies designed from the approximating discrete-time controlled Markov chain model, for the discounted/ergodic optimal control problems, in the true controlled diffusion model (as the sampling period approaches zero). To this end, we first construct suitable discrete-time controlled Markov chain models for which one can compute optimal policies and optimal values via several methods (such as value iteration, convex analytic method, reinforcement learning etc.). Then using a weak convergence technique, we show that the optimal policy designed for the discrete-time Markov chain model is near-optimal for the controlled diffusion model as the discrete-time model approaches the continuous-time model. This provides a practical approach for finding near-optimal control policies for controlled diffusions. Our conditions complement existing results in the literature, which have been arrived at via either probabilistic or PDE based methods.

math.OC

Nonzero-sum Discrete-time Stochastic Games with Risk-sensitive Ergodic Cost Criterion

In this paper we study infinite horizon nonzero-sum stochastic games for controlled discrete-time Markov chains on a Polish state space with risk-sensitive ergodic cost criterion. Under suitable assumptions we show that the associated ergodic optimality equations admit unique solutions. Finally, the existence of Nash-equilibrium in randomized stationary strategies is established by showing that an appropriate set-valued map has a fixed point.

math.OC

Near Optimality of Lipschitz and Smooth Policies in Controlled Diffusions

For optimal control of diffusions under several criteria, due to computational or analytical reasons, many studies have a apriori assumed control policies to be Lipschitz or smooth, often with no rigorous analysis on whether this restriction entails loss. While optimality of Markov/stationary Markov policies for expected finite horizon/infinite horizon (discounted/ergodic) cost and cost-up-to-exit time optimal control problems can be established under certain technical conditions, an optimal solution is typically only measurable in the state (and time, if the horizon is finite) with no apriori additional structural properties. In this paper, building on our recent work [S. Pradhan and S. Yüksel, Continuity of cost in Borkar control topology and implications on discrete space and time approximations for controlled diffusions under several criteria, Electronic Journal of Probability (2024)] establishing the regularity of optimal cost on the space of control policies under the Borkar control topology for a general class of diffusions, we establish near optimality of smooth/Lipschitz continuous policies for optimal control under expected finite horizon, infinite horizon discounted/average, and up-to-exit time cost criteria. Under mild assumptions, we first show that smooth/Lipschitz continuous policies are dense in the space of Markov/stationary Markov policies under the Borkar topology. Then utilizing the continuity of optimal costs as a function of policies on the space of Markov/stationary policies under the Borkar topology, we establish that optimal policies can be approximated by smooth/Lipschitz continuous policies with arbitrary precision. While our results are extensions of our recent work, the practical significance of an explicit statement and accessible presentation dedicated to Lipschitz and smooth policies, given their prominence in the literature, motivates our current paper.

math.OC

Robustness of Optimal Controlled Diffusions with Near-Brownian Noise via Rough Paths Theory

In this article we show a robustness theorem for controlled stochastic differential equations driven by approximations of Brownian motion. Often, Brownian motion is used as an idealized model of a diffusion where approximations such as Wong-Zakai, Karhnen-Loève or fractional Brownian motion are often seen as more physical. However, there has been extensive literature on solving control problems driven by Brownian motion and little on control problems driven by more realistic models that are only approximately Brownian. The question of robustness naturally arises from such approximations. We show robustness using rough paths theory, which allows for a pathwise theory of stochastic differential equations. To this end, in particular, we show that within the class of Lipschitz continuous control policies, an optimal solution for the Brownian idealized model is near optimal for a true system driven by a non-Brownian (but near-Brownian) noise.

math.OC

Robustness of Stochastic Optimal Control to Approximate Diffusion Models under Several Cost Evaluation Criteria

In control theory, typically a nominal model is assumed based on which an optimal control is designed and then applied to an actual (true) system. This gives rise to the problem of performance loss due to the mismatch between the true model and the assumed model. A robustness problem in this context is to show that the error due to the mismatch between a true model and an assumed model decreases to zero as the assumed model approaches the true model. We study this problem when the state dynamics of the system are governed by controlled diffusion processes. In particular, we will discuss continuity and robustness properties of finite horizon and infinite-horizon $α$-discounted/ergodic optimal control problems for a general class of non-degenerate controlled diffusion processes, as well as for optimal control up to an exit time. Under a general set of assumptions and a convergence criterion on the models, we first establish that the optimal value of the approximate model converges to the optimal value of the true model. We then establish that the error due to mismatch that occurs by application of a control policy, designed for an incorrectly estimated model, to a true model decreases to zero as the incorrect model approaches the true model. We will see that, compared to related results in the discrete-time setup, the continuous-time theory will let us utilize the strong regularity properties of solutions to optimality (HJB) equations, via the theory of uniformly elliptic PDEs, to arrive at strong continuity and robustness properties.

math.OC

Continuity of Cost in Borkar Control Topology and Implications on Discrete Space and Time Approximations for Controlled Diffusions under Several Criteria

We first show that the discounted cost, cost up to an exit time, and ergodic cost involving controlled non-degenerate diffusions are continuous on the space of stationary control policies when the policies are given a topology introduced by Borkar [V. S. Borkar, A topology for Markov controls, Applied Mathematics and Optimization 20 (1989), 55-62]. The same applies for finite horizon problems when the control policies are markov and the topology is revised to include time also as a parameter. We then establish that finite action/piecewise constant stationary policies are dense in the space of stationary Markov policies under this topology. Using the above mentioned continuity and denseness results we establish that finite action/piecewise constant policies approximate optimal stationary policies with arbitrary precision. This gives rise to the applicability of many numerical methods such as policy iteration and stochastic learning methods for discounted cost, cost up to an exit time, and ergodic cost optimal control problems in continuous-time. For the finite-horizon setup, we establish additionally near optimality of time-discretized policies by an analogous argument. We thus present a unified and concise approach for approximations directly applicable under several commonly adopted cost criteria.

math.OC

Ergodic Risk-Sensitive Control for Regime-Switching Diffusions

In this article, we study the ergodic risk-sensitive control problem for controlled regime-switching diffusions. Under a blanket stability hypothesis, we solve the associated nonlinear eigenvalue problem for weakly coupled systems and characterize the optimal stationary Markov controls via a suitable verification theorem. We also consider the near-monotone case and obtain the existence of principal eigenfunction and optimal stationary Markov controls.

math.OC

Nonzero-Sum Risk-Sensitive Stochastic Differential Games: A Multi-parameter Eigenvalue Problem Approach

We study nonzero-sum stochastic differential games with risk-sensitive ergodic cost criterion. Under certain conditions, using multi-parameter eigenvalue approach, we establish the existence of a Nash equilibrium in the space of stationary Markov strategies. We achieve our results by studying the relevant systems of coupled HJB equations. Exploiting the stochastic representation of the principal eigenfunctions we completely characterize Nash equilibrium points in the space of stationary Markov strategies.

math.OC

Discrete-time Zero-Sum Games for Markov chains with risk-sensitive average cost criterion

We study zero-sum stochastic games for controlled discrete time Markov chains with risk-sensitive average cost criterion with countable state space and Borel action spaces. The payoff function is nonnegative and possibly unbounded. Under a certain Lyapunov stability assumption on the dynamics, we establish the existence of a value and saddle point equilibrium. Further we completely characterize all possible saddle point strategies in the class of stationary Markov strategies. Finally, we present and analyze an illustrative example.

math.OC

Ergodic Risk-Sensitive Control of Markov Processes on Countable State Space Revisited

We consider a large family of discrete and continuous time controlled Markov processes and study an ergodic risk-sensitive minimization problem. Under a blanket stability assumption, we provide a complete analysis to this problem. In particular, we establish uniqueness of the value function and verification result for optimal stationary Markov controls, in addition to the existence results. We also revisit this problem under a near-monotonicity condition but without any stability hypothesis. Our results also include policy improvement algorithms both in discrete and continuous time frameworks.

math.OC

Zero-Sum Games for Continuous-time Markov Decision Processes with Risk-Sensitive Average Cost Criterion

We consider zero-sum stochastic games for continuous time Markov decision processes with risk-sensitive average cost criterion. Here the transition and cost rates may be unbounded. We prove the existence of the value of the game and a saddle-point equilibrium in the class of all stationary strategies under a Lyapunov stability condition. This is accomplished by establishing the existence of a principal eigenpair for the corresponding Hamilton-Jacobi-Isaacs (HJI) equation. This in turn is established by using the nonlinear version of Krein-Rutman theorem. We then obtain a characterization of the saddle-point equilibrium in terms of the corresponding HJI equation. Finally, we use a controlled population system to illustrate results.

math.OC

Nonzero-sum risk-sensitive continuous-time stochastic games with ergodic costs

We study nonzero-sum stochastic games for continuous time Markov decision processes on a denumerable state space with risk-sensitive ergodic cost criterion. Transition rates and cost rates are allowed to be unbounded. Under a Lyapunov type stability assumption, we show that the corresponding system of coupled HJB equations admits a solution which leads to the existence of a Nash equilibrium in stationary strategies. We establish this using an approach involving principal eigenvalues associated with the HJB equations. Furthermore, exploiting appropriate stochastic representation of principal eigenfunctions, we completely characterize Nash equilibria in the space of stationary Markov strategies.

math.OC

On the monotonicity property of the generalized eigenvalue for weakly-coupled cooperative elliptic systems

We consider general linear non-degenerate weakly-coupled cooperative elliptic systems and study certain monotonicity properties of the generalized principal eigenvalue in $\mathbb{R}^d$ with respect to the potential. It is shown that monotonicity on the right is equivalent to the recurrence property of the twisted operator which is, in turn, equivalent to the minimal growth property at infinity of the principal eigenfunctions. The strict monotonicity property of the principal eigenvalue is shown to be equivalent with the exponential stability of the twisted operators. An equivalence between the monotonicity property on the right and the stochastic representation of the principal eigenfunction is also established.

math.AP