Searcharxiv⌕ Search

arXiv subjects

Manuela L. Bujorianu

Publications and source records attributed to Manuela L. Bujorianu.

11 recordsLinked to original sources

Large Deviations in Safety-Critical Systems with Probabilistic Initial Conditions

We often rely on probabilistic measures -- e.g. event probability or expected time -- to characterize systems' safety. However, determining these quantities for extremely low-probability events is generally challenging, as standard safety methods usually struggle due to conservativeness, high-dimension scalability, tractability or numerical limitations. We address these issues by leveraging rigorous approximations grounded in the principles of Large Deviations theory. By assuming deterministic initial conditions, Large Deviations identifies a single dominant path in the low-noise limit as the most significant contributor to the rare-event probability: the instanton. We extend this result to incorporate stochastic uncertainty in the initial states, which is a common assumption in many applications. To that end, we determine an expression for the probability density of the initial states, conditioned on the unsafe rare event being observed. This expression gives access to the most probable initial conditions, as well as the most probable hitting time and path deviations, leading to the realization of the unsafe event. We demonstrate it's effectiveness by solving a high-dimensional and non-linear problem: a space collision.

math.OC↗

Provably Safe Reinforcement Learning for Stochastic Reach-Avoid Problems with Entropy Regularization

We consider the problem of learning the optimal policy for Markov decision processes with safety constraints. We formulate the problem in a reach-avoid setup. Our goal is to design online reinforcement learning algorithms that ensure safety constraints with arbitrarily high probability during the learning phase. To this end, we first propose an algorithm based on the optimism in the face of uncertainty (OFU) principle. Based on the first algorithm, we propose our main algorithm, which utilizes entropy regularization. We investigate the finite-sample analysis of both algorithms and derive their regret bounds. We demonstrate that the inclusion of entropy regularization improves the regret and drastically controls the episode-to-episode variability that is inherent in OFU-based safe RL algorithms.

cs.LG↗

Data-Driven Robust Safety Verification for Markov Decision Processes

In this paper, we propose a data-driven robust safety verification framework for stochastic dynamical systems modeled as Markov decision processes with time-varying and uncertain transition probabilities. Rather than assuming access to the exact nominal transition kernel, we consider the realistic setting where only samples from multiple system executions are available. These samples may correspond to different transition models inside an ambiguity set around the nominal transition kernel. Using these observations, we construct a unified ambiguity set that captures both inherent run-to-run variability in the transition dynamics and finite-sample statistical uncertainty. This ambiguity set is formalized through a Wasserstein-distance ball around a nominal empirical distribution and naturally induces an interval Markov decision process representation of the underlying system. Within this representation, we introduce a robust safety function that characterizes reach-avoid type probabilistic safety under all transition kernels consistent with the interval Markov decision process. We further derive high-confidence safety guarantees for the true, unknown time-varying system. A numerical example illustrates the applicability and effectiveness of the proposed approach.

eess.SY↗

An Online Multiobjective Policy Gradient for Long-run Average-reward Markov Decision Process

We propose a reinforcement learning (RL) framework for multi-objective decision-making, where the agent seeks to optimize a vector of rewards rather than a single scalar value. The objective is to ensure that the time-averaged reward vector converges asymptotically to a predefined target set. Since standard RL algorithms operate on scalar rewards, we introduce a dynamic scalarization mechanism guided by Blackwell's Approachability Theorem. This theorem enables adaptive updates of the scalarization vector to guarantee convergence toward the target set. Assuming ergodicity, the Markov chain induced by the learned policies admits a stationary distribution, ensuring all states recur with finite return times. Our algorithm exploits this property by defining an inner loop that applies a policy gradient method (with baseline) between successive visits to a designated recurrent state, enforcing Blackwell's condition at each iteration. An outer loop then updates the scalarization vector after each recurrence. We establish theoretical convergence of the long-run average reward vector to the target set and validate the approach through a numerical example.

eess.SY↗

Robust Correlated Equilibrium: Definition and Computation

We study N-player finite games with costs perturbed due to time-varying disturbances in the underlying system and to that end, we propose the concept of Robust Correlated Equilibrium that generalizes the definition of Correlated Equilibrium. Conditions under which the Robust Correlated Equilibrium exists are specified, and a decentralized algorithm for learning strategies that are optimal in the sense of Robust Correlated Equilibrium is proposed. The primary contribution of the paper is the convergence analysis of the algorithm and to that end, we propose a modification of the celebrated Blackwell's Approachability theorem to games with costs that are not just time-average, as in the original Blackwell's Approachability Theorem, but also include the time-average of previous algorithm iterates. The designed algorithm is applied to a practical water distribution network with pumps being the controllers and their costs being perturbed by uncertain consumption due to the consumers. Simulation results show that each controller achieves no regret, and empirical distributions converge to the Robust Correlated Equilibrium.

eess.SY↗

Safe Reinforcement Learning for Constrained Markov Decision Processes with Stochastic Stopping Time

In this paper, we present an online reinforcement learning algorithm for constrained Markov decision processes with a safety constraint. Despite the necessary attention of the scientific community, considering stochastic stopping time, the problem of learning optimal policy without violating safety constraints during the learning phase is yet to be addressed. To this end, we propose an algorithm based on linear programming that does not require a process model. We show that the learned policy is safe with high confidence. We also propose a method to compute a safe baseline policy, which is central in developing algorithms that do not violate the safety constraints. Finally, we provide simulation results to show the efficacy of the proposed algorithm. Further, we demonstrate that efficient exploration can be achieved by defining a subset of the state-space called proxy set.

cs.LG↗

Online Model-free Safety Verification for Markov Decision Processes Without Safety Violation

In this paper, we consider the problem of safety assessment for Markov decision processes without explicit knowledge of the model. We aim to learn probabilistic safety specifications associated with a given policy without compromising the safety of the process. To accomplish our goal, we characterize a subset of the state-space called proxy set, which contains the states that are near in a probabilistic sense to the forbidden set consisting of all unsafe states. We compute the safety function using the single-step temporal difference method. To this end, we relate the safety function computation to that of the value function estimation using temporal difference learning. Since the given control policy could be unsafe, we use a safe baseline subpolicy to generate data for learning. We then use an off-policy temporal difference learning method with importance sampling to learn the safety function corresponding to the given policy. Finally, we demonstrate our results using a numerical example

eess.SY↗

Safe Dynamic Programming

We incorporate safety specifications into dynamic programming. Explicitly, we address the minimization problem of a Markov decision process up to a stopping time with safety constraints. To incorporate safety into dynamic programming, we establish a formalism leaning upon the evolution equation. We show how to compute the safety function with the method of dynamic programming. In the last part of the paper, we develop several algorithms for safe dynamic programming.

math.OC↗

A new stochastic framework for ship capsizing

We present a new stochastic framework for studying ship capsize. It is a synthesis of two strands of transition state theory. The first is an extension of deterministic transition state theory to dissipative non-autonomous systems, together with a probability distribution over the forcing functions. The second is stochastic reachability and large deviation theory for transition paths in Markovian systems. In future work we aim to bring these together to make a tool for predicting capsize rate in different stochastic sea states, suggesting control strategies and improving designs.

math.DS↗

A coordination model for ultra-large scale systems of systems

The ultra large multi-agent systems are becoming increasingly popular due to quick decay of the individual production costs and the potential of speeding up the solving of complex problems. Examples include nano-robots, or systems of nano-satellites for dangerous meteorite detection, or cultures of stem cells for organ regeneration or nerve repair. The topics associated with these systems are usually dealt within the theories of intelligent swarms or biologically inspired computation systems. Stochastic models play an important role and they are based on various formulations of the mechanical statistics. In these cases, the main assumption is that the swarm elements have a simple behaviour and that some average properties can be deduced for the entire swarm. In contrast, complex systems in areas like aeronautics are formed by elements with sophisticated behaviour, which are even autonomous. In situations like this, a new approach to swarm coordination is necessary. We present a stochastic model where the swarm elements are communicating autonomous systems, the coordination is separated from the component autonomous activity and the entire swarm can be abstracted away as a piecewise deterministic Markov process, which constitutes one of the most popular model in stochastic control. Keywords: ultra large multi-agent systems, system of systems, autonomous systems, stochastic hybrid systems.

eess.SY↗

Proceedings Third International Workshop on Hybrid Autonomous Systems

The interest on autonomous systems is increasing both in industry and academia. Such systems must operate with limited human intervention in a changing environment and must be able to compensate for significant system failures without external intervention. The most appropriate models of autonomous systems can be found in the class of hybrid systems (which study continuous-state dynamic processes via discrete-state controllers) that interact with their environment. This workshop brings together researchers interested in all aspects of autonomy and resilience of hybrid systems.

eess.SY↗