SearcharxivSearch

arXiv subjects

Alexander Vladimirsky

Publications and source records attributed to Alexander Vladimirsky.

At least 19 recordsLinked to original sources

Partial health status observability and time horizon uncertainty in mean-field game epidemiological models

We introduce Mean-Field Game (MFG) epidemiological models, in which immunity either wanes with time in a fully observable way or disappears instantaneously with no direct observation (making a previously recovered individual fully susceptible again without realizing it). Both interpretations create computational challenges for rational noninfected individuals deciding on their contact rates based on their personal current immunity state and the changing epidemiological situation. Both require solving a forward-backward MFG system that includes PDEs (an advection-reaction equation for the immunity-structured population and a Hamilton-Jacobi-Bellman equation for the corresponding value function). We show how this can be done efficiently by solving a two-point boundary value problem for a system of approximating ODEs. We also show how the same approach can be extended to handle an initial uncertainty in the planning horizon.

math.OC

Behavioral patterns and mean-field games in epidemiological models

We introduce a new type of Mean Field Game epidemiological models, in which subpopulations have different behavioral patterns: some are viewed as "highly rational" (choosing Nash-equilibrium long-term strategies) while others follow pre-specified "non-rational" patterns (e.g., either sticking to their usual habits or trying to mimic those around them). Our model also allows for occasional behavioral switches, which rational individuals also take into account when formulating their Nash-equilibrium strategies. While this modeling approach is general, here we develop it for individuals choosing their "contact rates" within a particular Susceptible-Infected-Recovered-Susceptible-Dead (SIRSD) epidemics model. The latter is based on a frequency-based force of infection and the mortality rate that rapidly increases once the proportion of infected individuals exceeds some prescribed threshold, resulting in a strain on medical resources. Numerical tests illustrate the properties of our model and highlight the ways in which additional/non-rational behavioral patterns and behavioral switching increase the impact of infectious diseases. The paper aims to build a bridge between two distinct communities of epidemiological modelers and to promote the consideration of behavioral patterns in broader Mean Field Games literature.

q-bio.PE

Occasionally Observed Piecewise-deterministic Markov Processes

Piecewise-deterministic Markov processes (PDMPs) are often used to model abrupt changes in the global environment or capabilities of a controlled system. This is typically done by considering a set of "operating modes" (each with its own system dynamics and performance metrics) and assuming that the mode can switch stochastically while the system state evolves. Such models have a broad range of applications in engineering, economics, manufacturing, robotics, and biological sciences. Here, we introduce and analyze an "occasionally observed" version of mode-switching PDMPs. We show how such systems can be controlled optimally if the planner is not alerted to mode-switches as they occur but may instead have access to infrequent mode observations. We first develop a general framework for handling this through dynamic programming on a higher-dimensional mode-belief space. While quite general, this method is rarely practical due to the curse of dimensionality. We then discuss assumptions that allow for solving the same problem much more efficiently, with the computational costs growing linearly (rather than exponentially) with the number of modes. We use this approach to derive Hamilton-Jacobi-Bellman PDEs and quasi-variational inequalities encoding the optimal behavior for a variety of planning horizons (fixed, infinite, indefinite, random) and mode-observation schemes (at fixed times or on-demand). We discuss the computational challenges associated with each version and illustrate the resulting methods on test problems from surveillance-evading path planning. We also include an example based on robotic navigation: a Mars rover that minimizes the expected time to target while accounting for the possibility of unobserved/incremental damages and dynamics-altering breakdowns.

math.OC

Monotone Causality in Opportunistically Stochastic Shortest Path Problems

When traveling through a graph with an accessible deterministic path to a target, is it ever preferable to resort to stochastic node-to-node transitions instead? And if so, what are the conditions guaranteeing that such a stochastic optimal routing policy can be computed efficiently? We aim to answer these questions here by defining a class of Opportunistically Stochastic Shortest Path (OSSP) problems and deriving sufficient conditions for applicability of non-iterative label-setting methods. The usefulness of this framework is demonstrated in two very different contexts: numerical analysis and autonomous vehicle routing. We use OSSPs to derive causality conditions for semi-Lagrangian discretizations of anisotropic Hamilton-Jacobi equations. We also use a Dijkstra-like method to solve OSSPs optimizing the timing and urgency of lane change maneuvers for an autonomous vehicle navigating road networks with a heterogeneous traffic load.

math.OC

Optimality and robustness in path-planning under initial uncertainty

Classical deterministic optimal control problems assume full information about the controlled process. The theory of control for general partially-observable processes is powerful, but the methods are computationally expensive and typically address the problems with stochastic dynamics and continuous (directly unobserved) stochastic perturbations. In this paper we focus on path planning problems which are in between -- deterministic, but with an initial uncertainty on either the target or the running cost on parts of the domain. That uncertainty is later removed at some time $T$, and the goal is to choose the optimal trajectory until then. We address this challenge for three different models of information acquisition: with fixed $T$, discretely distributed and exponentially distributed random $T$. We develop models and numerical methods suitable for multiple notions of optimality: based on the average-case performance, the worst-case performance, the average constrained by the worst, the average performance with probabilistic constraints on the bad outcomes, risk-sensitivity, and distributional-robustness. We illustrate our approach using examples of pursuing random targets identified at a (possibly random) later time $T$.

math.OC

Risk-aware stochastic control of a sailboat

Sailboat path-planning is a natural hybrid control problem (due to continuous steering and occasional "tack-switching" maneuvers), with the actual path-to-target greatly affected by stochastically evolving wind conditions. Previous studies have focused on finding risk-neutral policies that minimize the expected time of arrival. In contrast, we present a robust control approach, which maximizes the probability of arriving before a specified deadline/threshold. Our numerical method recovers the optimal risk-aware (and threshold-specific) policies for all initial sailboat positions and a broad range of thresholds simultaneously. This is accomplished by solving two quasi-variational inequalities based on second-order Hamilton-Jacobi-Bellman (HJB) PDEs with degenerate parabolicity. Monte-Carlo simulations show that risk-awareness in sailing is particularly useful when a carefully calculated bet on the evolving wind direction might yield a reduction in the number of tack-switches.

math.OC

Quantifying and managing uncertainty in piecewise-deterministic Markov processes

In piecewise-deterministic Markov processes (PDMPs) the state of a finite-dimensional system evolves continuously, but the evolutive equation may change randomly as a result of discrete switches. A running cost is integrated along the corresponding piecewise-deterministic trajectory up to the termination to produce the cumulative cost of the process. We address three natural questions related to uncertainty in cumulative cost of PDMP models: (1) how to compute the Cumulative Distribution Function (CDF) of the cumulative cost when the switching rates are fully known; (2) how to accurately bound the CDF when the switching rates are uncertain; and (3) assuming the PDMP is controlled, how to select a control to optimize that CDF. In all three cases, our approach requires posing a system of suitable hyperbolic partial differential equations, which are then solved numerically on an augmented state space. We illustrate our method using simple examples of trajectory planning under uncertainty for several 1D and 2D first-exit time problems. In the Appendix, we also apply this method to a model of fish harvesting in an environment with random switches in carrying capacity.

math.OC

Surveillance Evasion Through Bayesian Reinforcement Learning

We consider a task of surveillance-evading path-planning in a continuous setting. An Evader strives to escape from a 2D domain while minimizing the risk of detection (and immediate capture). The probability of detection is path-dependent and determined by the spatially inhomogeneous surveillance intensity, which is fixed but a priori unknown and gradually learned in the multi-episodic setting. We introduce a Bayesian reinforcement learning algorithm that relies on a Gaussian Process regression (to model the surveillance intensity function based on the information from prior episodes), numerical methods for Hamilton-Jacobi PDEs (to plan the best continuous trajectories based on the current model), and Confidence Bounds (to balance the exploration vs exploitation). We use numerical experiments and regret metrics to highlight the significant advantages of our approach compared to traditional graph-based algorithms of reinforcement learning.

cs.LG

Optimal Path-Planning with Random Breakdowns

We propose a model for path-planning based on a single performance metric that accurately accounts for the the potential (spatially inhomogeneous) cost of breakdowns and repairs. These random breakdowns (or system faults) happen at a known, spatially inhomogeneous rate. Our model includes breakdowns of two types: total, which halt all movement until an in-place repair is completed, and partial, after which movement continues in a damaged state toward a repair depot. We use the framework of piecewise-deterministic Markov processes to describe the optimal policy for all starting locations. We also introduce an efficient numerical method that uses hybrid value-policy iterations to solve the resulting system of Hamilton-Jacobi-Bellman PDEs. Our method is illustrated through a series of computational experiments that highlight the dependence of optimal policies on the rate and type of breakdowns, with one of them based on Martian terrain data near Jezero Crater.

math.OC

Optimal Driving Under Traffic Signal Uncertainty

We study driver's optimal trajectory planning under uncertainty in the duration of a traffic light's green phase. We interpret this as an optimal control problem with an objective of minimizing the expected cost based on the fuel use, discomfort from rapid velocity changes, and time to destination. Treating this in the framework of dynamic programming, we show that the probability distribution on green phase durations gives rise to a sequence of Hamilton-Jacobi-Bellman PDEs, which are then solved numerically to obtain optimal acceleration/braking policy in feedback form. Our numerical examples illustrate the approach and highlight the role of conflicting goals and uncertainty in shaping drivers' behavior.

math.OC

Stochastic Optimal Control of a Sailboat

In match race sailing, competitors must steer their boats upwind in the presence of unpredictably evolving weather. Combined with the tacking motion necessary to make upwind progress, this makes it natural to model their path-planning as a hybrid stochastic optimal control problem. Dynamic programming provides the tools for solving these, but the computational cost can be significant. We greatly accelerate a semi-Lagrangian iterative approach of Ferretti and Festa (R. Ferretti and A. Festa, "Optimal Route Planning for Sailing Boats: A Hybrid Formulation", J Optim Theory Appl (2019)) by reducing the state space dimension and designing an adaptive timestep discretization that is very nearly causal. We also provide a more accurate tack-switching operator by integrating over potential wind states after the switch. The method is illustrated through a series of simulations with varying stochastic wind conditions.

math.OC

Control-Theoretic Models of Environmental Crime

We present two models of perpetrators' decision-making in extracting resources from a protected area. It is assumed that the authorities conduct surveillance to counter the extraction activities, and that perpetrators choose their post-extraction paths to balance the time/hardship of travel against the expected losses from a possible detection. In our first model, the authorities are assumed to use ground patrols and the protected resources are confiscated as soon as the extractor is observed with them. The perpetrators' path-planning is modeled using the optimal control of randomly-terminated process. In our second model, the authorities use aerial patrols, with the apprehension of perpetrators and confiscation of resources delayed until their exit from the protected area. In this case the path-planning is based on multi-objective dynamic programming. Our efficient numerical methods are illustrated on several examples with complicated geometry and terrain of protected areas, non-uniform distribution of protected resources, and spatially non-uniform detection rates due to aerial or ground patrols.

math.OC

Time-Dependent Surveillance-Evasion Games

Surveillance-Evasion (SE) games form an important class of adversarial trajectory-planning problems. We consider time-dependent SE games, in which an Evader is trying to reach its target while minimizing the cumulative exposure to a moving enemy Observer. That Observer is simultaneously aiming to maximize the same exposure by choosing how often to use each of its predefined patrol trajectories. Following the framework introduced in Gilles and Vladimirsky (arXiv:1812.10620), we develop efficient algorithms for finding Nash Equilibrium policies for both players by blending techniques from semi-infinite game theory, convex optimization, and multi-objective dynamic programming on continuous planning spaces. We illustrate our method on several examples with Observers using omnidirectional and angle-restricted sensors on a domain with occluding obstacles.

math.OC

Evasive path planning under surveillance uncertainty

The classical setting of optimal control theory assumes full knowledge of the process dynamics and the costs associated with every control strategy. The problem becomes much harder if the controller only knows a finite set of possible running cost functions, but has no way of checking which of these running costs is actually in place. In this paper we address this challenge for a class of evasive path planning problems on a continuous domain, in which an Evader needs to reach a target while minimizing his exposure to an enemy Observer, who is in turn selecting from a finite set of known surveillance plans. Our key assumption is that both the evader and the observer need to commit to their (possibly probabilistic) strategies in advance and cannot immediately change their actions based on any newly discovered information about the opponent's current position. We consider two types of evader behavior: in the first one, a completely risk-averse evader seeks a trajectory minimizing his {\em worst-case} cumulative observability, and in the second, the evader is concerned with minimizing the {\em average-case} cumulative observability. The latter version is naturally interpreted as a semi-infinite strategic game, and we provide an efficient method for approximating its Nash equilibrium. The proposed approach draws on methods from game theory, convex optimization, optimal control, and multiobjective dynamic programming. We illustrate our algorithm using numerical examples and discuss the computational complexity, including for the generalized version with multiple evaders.

math.OC

Optimizing adaptive cancer therapy: dynamic programming and evolutionary game theory

Recent clinical trials have shown that the adaptive drug therapy can be more efficient than a standard MTD-based policy in treatment of cancer patients. The adaptive therapy paradigm is not based on a preset schedule; instead, the doses are administered based on the current state of tumor. But the adaptive treatment policies examined so far have been largely ad hoc. In this paper we propose a method for systematically optimizing the rules of adaptive policies based on an Evolutionary Game Theory model of cancer dynamics. Given a set of treatment objectives, we use the framework of dynamic programming to find the optimal treatment strategies. In particular, we optimize the total drug usage and time to recovery by solving a Hamilton-Jacobi-Bellman equation based on a mathematical model of tumor evolution. We compare adaptive/optimal treatment strategy with MTD-based treatment policy. We show that optimal treatment strategies can dramatically decrease the total amount of drugs prescribed as well as increase the fraction of initial tumour states from which the recovery is possible. We also examine the optimization trade-offs between the total administered drugs and recovery time. The adaptive therapy combined with optimal control theory is a promising concept in the cancer treatment and should be integrated into clinical trial design.

q-bio.QM

Anisotropic Challenges in Pedestrian Flow Modeling

Macroscopic models of crowd flow incorporating individual pedestrian choices present many analytic and computational challenges. Anisotropic interactions are particularly subtle, both in terms of describing the correct "optimal" direction field for the pedestrians and ensuring that this field is uniquely defined. We develop sufficient conditions, which establish a range of "safe" densities and parameter values for each model. We illustrate our approach by analyzing several established intra-crowd and inter-crowd models. For the two-crowd case, we also develop sufficient conditions for the uniqueness of Nash Equilibria in the resulting non-zero-sum game.

math.OC

Corner cases, singularities, and dynamic factoring

In Eikonal equations, rarefaction is a common phenomenon known to degrade the rate of convergence of numerical methods. The `factoring' approach alleviates this difficulty by deriving a PDE for a new (locally smooth) variable while capturing the rarefaction-related singularity in a known (non-smooth) `factor'. Previously this technique was successfully used to address rarefaction fans arising at point sources. In this paper we show how similar ideas can be used to factor the 2D rarefactions arising due to nonsmoothness of domain boundaries or discontinuities in PDE coefficients. Locations and orientations of such rarefaction fans are not known in advance and we construct a `just-in-time factoring' method that identifies them dynamically. The resulting algorithm is a generalization of the Fast Marching Method originally introduced for the regular (unfactored) Eikonal equations. We show that our approach restores the first-order convergence and illustrate it using a range of maze navigation examples with non-permeable and `slowly permeable' obstacles.

math.NA

Piecewise-Deterministic Optimal Path Planning

We consider piecewise-deterministic optimal control problems in which the environment randomly switches among several deterministic modes, and the goal is to optimize the expected cost up to the termination while taking the likelihood of future mode-switches into account. The dynamic programming approach yields a weakly-coupled system of Hamilton-Jacobi-Bellman PDEs satisfied by the value functions of individual modes. We derive and implement semi-Lagrangian and Eulerian numerical schemes for that system. We further recover simpler, "asymptotically optimal" controls and test their performance in non-asymptotic regimes. Our approach is illustrated on a simple anisotropic path-planning problem: the time-optimal control for a boat affected by randomly switching winds.

math.OC