SearcharxivSearch

arXiv subjects

Sumith Reddy Anugu

Publications and source records attributed to Sumith Reddy Anugu.

6 recordsLinked to original sources

Exponential rate of convergence of relative value iteration algorithms for ergodic controls of diffusions

In this paper, we investigate the rate of convergence of the relative value iteration (RVI) algorithms for diffusions in $\mathbb{R}^d$ under both the conventional ergodic cost (CEC) and ergodic risk-sensitive cost (ERSC) criteria, and under the uniform exponential stability condition. The existing RVI algorithms for the CEC and ERSC problems solve the associated initial value Hamilton-Jacobi-Bellman type equations whose solutions are shown to converge asymptotically to the corresponding optimal values. However, the rates of convergence for such algorithms have remained open. This paper proposes discrete-time implementations for the RVI algorithms based on slight modifications of the associated PDEs, and proves that the rates of convergence of these RVI algorithms are exponential under a weighted sup-norm. These implementations have discrete-time iterates that can be explicitly expressed as recursive systems. The difference between these iterates and the desired value function in the CEC case can then be expressed in terms of the associated Markov kernels. Similarly, this can be done for the logarithms of the corresponding iterates and desired value function in the ERSC case in terms of the associated Markov kernels for the extended diffusion. As a result, we are able to prove the desirable contraction properties in order to establish the exponential rate of convergence by making use of a weighted semi-norm in which Markov kernel acts a contraction.

math.OC

Jacobi-like relative value iteration algorithms for ergodic risk-sensitive control of Markov chains

We propose a Jacobi-like relative value iteration (RVI) algorithm and a Gauss-Seidel-like implementation for the ergodic risk-sensitive control (ERSC) problem of a controlled discrete time Markov chain (DTMC) on a finite state space. Under the assumption that the DTMC is irreducible and recurrent under every stationary Markov policy, we prove that the iterates of the proposed RVI algorithms converge at a geometric rate. The main challenge stems from the multiplicative structure of the ERSC cost criterion and the associated Bellman-like operators, which prevents us from adapting the analogous global contraction and bi-Lipschitz continuity properties that underlie the proof of convergence in the average cost setting. We overcome this by establishing local contraction properties for the risk-sensitive Bellman-like operators and a local bi-Lipschitz continuity property for their fixed points, and use these properties to show the iterates converge geometrically. We conclude by implementing our proposed RVI algorithms on two examples: service effort control for a single-server queue of finite capacity, and maximizing the exit rate from a finite domain (on a graph).

math.OC

Small noise asymptotics for a class of jump-diffusions with heavy tails for large times

In this work, we investigate positive recurrent Lévy diffusions driven by appropriately scaled Brownian motion and $α$-stable process (with $1<α<2$) in the small noise regime. Supposing that in the vanishing noise limit, our Lévy diffusion approaches a deterministic system with a unique asymptotically stable fixed point, we show that the limiting behavior of the one-dimensional marginal distribution at large times is dictated by the optimal value of a deterministic control problem, just as in the classical case of diffusions driven by small variance Brownian motion. In our case, the control is allowed to have two parts: continuous control and impulse control.

math.PR

Ergodic Risk Sensitive Control of Diffusions under a General Structural Hypothesis

We study the infinite-horizon average (ergodic) risk sensitive control problem for diffusion processes under a general structural hypothesis: there is a partition of state space into two subsets, where the controlled diffusion process satisfies a Foster-Lyapunov type drift condition in one subset, under any stationary Markov control, while the near-monotonicity condition is satisfied with the running cost function being inf-compact in its complement. Under these conditions, we completely characterize the optimal stationary Markov controls. To prove this, we consider an inf-compact perturbation to the running cost over the entire space such that the resulting ergodic risk sensitive control problem is well-defined and then use the corresponding existing results. The heart of the analysis lies in exploiting the variational formula of exponential functionals of Brownian motion and applying it to the objective exponential cost function of the controlled diffusion. This representation facilitates us to view the risk sensitive cost for any stationary Markov control as the optimal value of a control problem of an extended diffusion involving a new auxiliary control where the optimal criterion is to maximize the associated long-run average cost criterion that is a difference of the original running cost and an extra term that is quadratic in the auxiliary control. The main difficulty in using this approach lies in the fact that tightness of mean empirical measures of the extended diffusion is not a priori implied by the analogous tightness property of the original diffusion. We overcome this by establishing a priori estimates for the extended diffusion associated with the nearly optimal auxiliary controls.

math.OC

Strong and weak quantitative estimates in slow-fast diffusions using filtering techniques

The behavior of slow-fast diffusions as the separation of scale diverges is a well-studied problem in the literature. In this short paper, we revisit this problem and obtain a new proof of existing strong quantitative convergence estimates (in particular, $L^2$ estimates) and weak convergence estimates in terms of $n$ (the parameter associated with the separation of scales). In particular, we obtain the rate of $n^{-\frac{1}{2}}$ in the strong convergence estimates and the rate of $n^{-1}$ for weak convergence estimate which are already known to be optimal in the literature. We achieve this using nonlinear filtering theory where we represent the evolution of fast diffusion in terms of its conditional distribution given the slow diffusion. We then use the well-known Kushner-Stratanovich equation which gives the evolution of the conditional distribution of the fast diffusion given the slow diffusion and establish that this conditional distribution approaches the invariant measure of the ``frozen" diffusion (obtained by freezing the slow variable in the evolution equation of the fast diffusion). At the heart of the analysis lies a key estimate of a weighted Lipschitz distance like function between a generic one-parameter family of measures and the family of unique invariant measures (of the ``frozen" diffusion parametrized by a path). This estimate is in terms of the operator norm of the dual of the infinitesimal generator of the ``frozen" diffusion.

math.OC

Ergodic Risk Sensitive Control of Markovian Multiclass Many-Server Queues with Abandonment

We study the optimal scheduling problem for a Markovian multiclass queueing network with abandonment in the Halfin--Whitt regime, under the long run average (ergodic) risk sensitive cost criterion. The objective is to prove asymptotic optimality for the optimal control arising from the corresponding ergodic risk sensitive control (ERSC) problem for the limiting diffusion. In particular, we show that the optimal ERSC value associated with the diffusion-scaled queueing process converges to that of the limiting diffusion in the asymptotic regime. The challenge that ERSC poses is that one cannot express the ERSC cost as an expectation over the mean empirical measure associated with the queueing process, unlike in the usual case of a long run average (ergodic) cost. We develop a novel approach by exploiting the variational representations of the limiting diffusion and the Poisson-driven queueing dynamics, which both involve certain auxiliary controls. The ERSC costs for both the diffusion-scaled queueing process and the limiting diffusion can be represented as the integrals of an extended running cost over a mean empirical measure associated with the corresponding extended processes using these auxiliary controls. For the lower bound proof, we exploit the connections of the ERSC problem for the limiting diffusion with a two-person zero-sum stochastic differential game. We also make use of the mean empirical measures associated with the extended limiting diffusion and diffusion-scaled processes with the auxiliary controls. One major technical challenge in both lower and upper bound proofs, is to establish the tightness of the aforementioned mean empirical measures for the extended processes. We identify nearly optimal controls appropriately in both cases so that the existing ergodicity properties of the limiting diffusion and diffusion-scaled queueing processes can be used.

math.PR