SearcharxivSearch

arXiv subjects

Abhijit Mazumdar

Publications and source records attributed to Abhijit Mazumdar.

7 recordsLinked to original sources

Provably Safe Reinforcement Learning for Stochastic Reach-Avoid Problems with Entropy Regularization

We consider the problem of learning the optimal policy for Markov decision processes with safety constraints. We formulate the problem in a reach-avoid setup. Our goal is to design online reinforcement learning algorithms that ensure safety constraints with arbitrarily high probability during the learning phase. To this end, we first propose an algorithm based on the optimism in the face of uncertainty (OFU) principle. Based on the first algorithm, we propose our main algorithm, which utilizes entropy regularization. We investigate the finite-sample analysis of both algorithms and derive their regret bounds. We demonstrate that the inclusion of entropy regularization improves the regret and drastically controls the episode-to-episode variability that is inherent in OFU-based safe RL algorithms.

cs.LG

Data-Driven Robust Safety Verification for Markov Decision Processes

In this paper, we propose a data-driven robust safety verification framework for stochastic dynamical systems modeled as Markov decision processes with time-varying and uncertain transition probabilities. Rather than assuming access to the exact nominal transition kernel, we consider the realistic setting where only samples from multiple system executions are available. These samples may correspond to different transition models inside an ambiguity set around the nominal transition kernel. Using these observations, we construct a unified ambiguity set that captures both inherent run-to-run variability in the transition dynamics and finite-sample statistical uncertainty. This ambiguity set is formalized through a Wasserstein-distance ball around a nominal empirical distribution and naturally induces an interval Markov decision process representation of the underlying system. Within this representation, we introduce a robust safety function that characterizes reach-avoid type probabilistic safety under all transition kernels consistent with the interval Markov decision process. We further derive high-confidence safety guarantees for the true, unknown time-varying system. A numerical example illustrates the applicability and effectiveness of the proposed approach.

eess.SY

Distributionally Robust Safety Verification for Markov Decision Processes

In this paper, we propose a distributionally robust safety verification method for Markov decision processes where only an ambiguous transition kernel is available instead of the precise transition kernel. We define the ambiguity set around the nominal distribution by considering a Wasserstein distance. To this end, we introduce a robust safety function to characterize probabilistic safety in the face of uncertain transition probability. First, we obtain an upper bound on the robust safety function in terms of a distributionally robust Q-function. Then, we present a convex program-based distributionally robust Q-iteration algorithm to compute the robust Q-function. By considering a numerical example, we demonstrate our theoretical results.

eess.SY

Safe Reinforcement Learning for Constrained Markov Decision Processes with Stochastic Stopping Time

In this paper, we present an online reinforcement learning algorithm for constrained Markov decision processes with a safety constraint. Despite the necessary attention of the scientific community, considering stochastic stopping time, the problem of learning optimal policy without violating safety constraints during the learning phase is yet to be addressed. To this end, we propose an algorithm based on linear programming that does not require a process model. We show that the learned policy is safe with high confidence. We also propose a method to compute a safe baseline policy, which is central in developing algorithms that do not violate the safety constraints. Finally, we provide simulation results to show the efficacy of the proposed algorithm. Further, we demonstrate that efficient exploration can be achieved by defining a subset of the state-space called proxy set.

cs.LG

Online Model-free Safety Verification for Markov Decision Processes Without Safety Violation

In this paper, we consider the problem of safety assessment for Markov decision processes without explicit knowledge of the model. We aim to learn probabilistic safety specifications associated with a given policy without compromising the safety of the process. To accomplish our goal, we characterize a subset of the state-space called proxy set, which contains the states that are near in a probabilistic sense to the forbidden set consisting of all unsafe states. We compute the safety function using the single-step temporal difference method. To this end, we relate the safety function computation to that of the value function estimation using temporal difference learning. Since the given control policy could be unsafe, we use a safe baseline subpolicy to generate data for learning. We then use an off-policy temporal difference learning method with importance sampling to learn the safety function corresponding to the given policy. Finally, we demonstrate our results using a numerical example

eess.SY

On Passivity, Feedback Passivity, And Feedback Passivity Over Erasure Network: A Piecewise Affine Approximation Approach

In this paper, we deal with the problem of passivity and feedback passification of smooth discrete-time nonlinear systems by considering their piecewise affine approximations. Sufficient conditions are derived for passivity and feedback passivity. These results are then extended to systems that operate over Gilbert-Elliott type communication channels. As a special case, results for feedback passivity of piecewise affine systems over a lossy channel are also derived.

eess.SY

$H_{\infty}$ Optimal Control of Jump Systems Over Multiple Lossy Communication Channels

In this paper, we consider the $H_{\infty}$ optimal control problem for a Markovian jump linear system (MJLS) over a lossy communication network. It is assumed that the controller communicates with each actuator through a different communication channel. We solve the $H_{\infty}$ optimization problem for a Transmission Control Protocol (TCP) using the theory of dynamic games and obtain a state-feedback controller. The infinite horizon $H_{\infty}$ optimization problem is analyzed as a limiting case of the finite horizon optimization problem. Then, we obtain the corresponding state-feedback controller, and show that it stabilizes the closed-loop system in the face of random packet dropouts.

eess.SY