SearcharxivSearch

arXiv subjects

Alexey Piunovskiy

Publications and source records attributed to Alexey Piunovskiy.

15 recordsLinked to original sources

Threshold Structure of Optimal Policies in Restart POMDPs

We study a Restart POMDP (Partially Observable Markov Decision Process) on a general Borel state space, where the controller either lets the hidden state evolve unobserved or restarts the system and observes the new state. Exploiting a sufficient-statistic representation consisting of the last observed state and the elapsed time since restart, we reduce the problem to a fully observed MDP. Under a natural one-step cost deterioration condition, we prove that optimal policies have a threshold structure in the elapsed time for both the discounted and total undiscounted cost criteria. When the state space is partially ordered and the kernel is stochastically monotone, we further show that the optimal threshold is nonincreasing in the state. For the average cost criterion, under additional assumptions of geometric ergodicity and domination of the transient gain, we establish analogous threshold results via the vanishing discount approach, after showing the uniform boundedness of the optimal thresholds and relative value functions.

math.OC

Primal-dual programs for the constrained optimal impulse control: discounted model

This paper studies constrained optimal impulse control problems of a deterministic system described by a (semi)flow, where the performance measures are the discounted total costs including both the costs incurred with applying impulses as well as running costs. We formulate the relaxed problem and the associated primal convex programs in measures together with the dual programs, and establish the relevant duality results. As an application, we formulate and justify a general procedure for obtaining optimal $(J+1)$-mixed strategies for the original impulse control problems. This procedure is illustrated with a solved example.

math.OC

On the continuity of the projection mapping from strategic measures to occupation measures in absorbing Markov decision processes

In this paper, we prove the following assertion for an absorbing Markov decision process (MDP) with the given initial distribution, which is also assumed to be semi-continuous: the continuity of the projection mapping from the space of strategic measures to the space of occupation measures, both endowed with their weak topologies, is equivalent to the MDP model being uniformly absorbing. An example demonstrates, among other interesting scenarios, that for an absorbing (but not uniformly absorbing) semi-continuous MDP with the given initial distribution, the space of occupation measures can fail to be compact in the weak topology.

math.OC

On gradual-impulse control of continuous-time Markov decision processes with exponential utility

In this paper, we consider the gradual-impulse control problem of continuous-time Markov decision processes, where the system performance is measured by the expectation of the exponential utility of the total cost. We prove, under very general conditions on the system primitives, the existence of a deterministic stationary optimal policy out of a more general class of policies. Policies that we consider allow multiple simultaneous impulses, randomized selection of impulses with random effects, relaxed gradual controls, and accumulation of jumps. After characterizing the value function using the optimality equation, we reduce the continuous-time gradual-impulse control problem to an equivalent simple discrete-time Markov decision process, whose action space is the union of the sets of gradual and impulsive actions.

math.OC

Extreme occupation measures in Markov decision processes with a cemetery

In this paper, we consider a Markov decision process (MDP) with a Borel state space $\textbf{X}\cup\{Δ\}$, where $Δ$ is an absorbing state (cemetery), and a Borel action space $\textbf{A}$. We consider the space of finite occupation measures restricted on $\textbf{X}\times \textbf{A}$, and the extreme points in it. It is possible that some strategies have infinite occupation measures. Nevertheless, we prove that every finite extreme occupation measure is generated by a deterministic stationary strategy. Then, for this MDP, we consider a constrained problem with total undiscounted criteria and $J$ constraints, where the cost functions are nonnegative. By assumption, the strategies inducing infinite occupation measures are not optimal. Then, our second main result is that, under mild conditions, the solution to this constrained MDP is given by a mixture of no more than $J+1$ occupation measures generated by deterministic stationary strategies.

math.OC

On the structure of optimal solutions in a mathematical programming problem in a convex space

We consider an optimization problem in a convex space $E$ with an affine objective function, subject to $J$ constraints in the forms of inequalities on some other affine functions, where $J$ is a given nonnegative integer. Under suitable conditions, we apply the Feinberg-Shwartz lemma in finite dimensional convex analysis to show that there exists an optimal solution, which is in the form of a mixture of no more than $J+1$ extreme points of $E$. It seems that in the current setup, this result has not yet been made available, because the concerned problem does not fit into the framework of standard convex optimization problems.

math.OC

Gradual-impulsive control for continuous-time Markov decision processes with total undiscounted costs and constraints: linear programming approach via a reduction method

We consider the constrained optimal control problem for the gradual-impulsive CTMDP model with the performance criteria being the expected total undiscounted costs (from the running cost and the cost from each time an impulse being applied). The discounted model is covered as a special case. We justify fully a reduction method, and close an open issue in the previous literature. The reduction method induces an equivalent but simpler standard CTMDP model with gradual control only, based on which, we establish effectively, under rather natural conditions, a linear programming approach for solving the concerned constrained optimal control problem.

math.OC

Modelling Ethnogenesis

Following the ideas of L.N.Gumilev, we introduce the mathematical model of ethnogenesis which describes the dynamics of subgroups in the developing polity in terms of ordinary differential equations. The bust dynamics associated with the rise and fall of civilisations is modelled as an excitation process, which is the non-linear phenomenon, well known in mathematical biology. We consider deterministic as well as the stochastic version of the model. We also expand the model to study the interaction between two polities undergoing ethnogenesis. Investigation is performed using analytical methods as well as numerical integration (i.e. MATLAB simulation).

math.DS

Turnpikes and Random Walk

In this paper we revise the theory of turnpikes in discounted Markov decision processes, prove the turnpike theorem for the undiscounted model and apply the results to the specific random walk.

math.PR

Linear programming approach to optimal impulse control problems with functional constraints

This paper considers an optimal impulse control problem of dynamical systems generated by a flow. The performance criteria are total costs over the infinite time horizon. Apart from the main performance to be minimized, there are multiple constraints on performance functionals of a similar type. Under a natural set of compactness-continuity conditions on the system primitives, we establish a linear programming approach, and prove the existence of a stationary optimal control strategy out of a more general class of randomized strategies. This is done by making use of the tools from Markov decision processes.

math.OC

Aggregated occupation measures and linear programming approach to constrained impulse control problems

For a constrained optimal impulse control problem of an abstract dynamical system, we introduce the occupation measures along with aggregated occupation measures and present two associated linear programs. We prove that the two linear programs are equivalent under appropriate conditions, and each linear program gives rise to an optimal strategy in the original impulse control problem.

math.OC

Optimal Impulse Control of Dynamical Systems

Using the tools of the Markov Decision Processes, we justify the dynamic programming approach to the optimal impulse control of deterministic dynamical systems. We prove the equivalence of the integral and differential forms of the optimality equation. The theory is illustrated by an example from mathematical epidemiology. The developed methods can be also useful for the study of piecewise deterministic Markov processes.

math.OC

Hitting Times in Markov Chains with Restart and their Application to Network Centrality

Motivated by applications in telecommunications, computer scienceand physics, we consider a discrete-time Markov process withrestart. At each step the process eitherwith a positive probability restarts from a given distribution, orwith the complementary probability continues according to a Markovtransition kernel. The main contribution of the present work is thatwe obtain an explicit expression for the expectation of the hittingtime (to a given target set) of the process with restart.The formula is convenient when considering the problem of optimizationof the expected hitting time with respect to the restart probability.We illustrate our results with two examplesin uncountable and countable state spaces andwith an application to network centrality.

cs.PF

Note on discounted continuous-time Markov decision processes with a lower bounding function

In this paper, we consider the discounted continuous-time Markov decision process (CTMDP) with a lower bounding function. In this model, the negative part of each cost rate is bounded by the drift function, say $w$, whereas the positive part is allowed to be arbitrarily unbounded. Our focus is on the existence of a stationary optimal policy for the discounted CTMDP problems out of the more general class. Both constrained and unconstrained problems are considered. Our investigations are based on a useful transformation for nonhomogeneous Markov pure jump processes that has not yet been widely applied to the study of CTMDPs. This technique was not employed in previous literature, but it clarifies the roles of the imposed conditions in a rather transparent way. As a consequence, we withdraw and weaken several conditions commonly imposed in the literature.

math.OC

Discounted Continuous-time Markov Decision Processes with Unbounded Rates: the Dynamic Programming Approach

This paper deals with unconstrained discounted continuous-time Markov decision processes in Borel state and action spaces. Under some conditions imposed on the primitives, allowing unbounded transition rates and unbounded (from both above and below) cost rates, we show the regularity of the controlled process, which ensures the underlying models to be well defined. Then we develop the dynamic programming approach by showing that the Bellman equation is satisfied (by the optimal value). Finally, under some compactness-continuity conditions, we obtain the existence of a deterministic stationary optimal policy out of the class of randomized history-dependent policies.

math.OC