SearcharxivSearch

arXiv subjects

Enric Ribera Borrell

Publications and source records attributed to Enric Ribera Borrell.

5 recordsLinked to original sources

Reinforcement Learning with Random Time Horizons

We extend the standard reinforcement learning framework to random time horizons. While the classical setting typically assumes finite and deterministic or infinite runtimes of trajectories, we argue that multiple real-world applications naturally exhibit random (potentially trajectory-dependent) stopping times. Since those stopping times typically depend on the policy, their randomness has an effect on policy gradient formulas, which we (mostly for the first time) derive rigorously in this work both for stochastic and deterministic policies. We present two complementary perspectives, trajectory or state-space based, and establish connections to optimal control theory. Our numerical experiments demonstrate that using the proposed formulas can significantly improve optimization convergence compared to traditional approaches.

cs.LG

Connecting Stochastic Optimal Control and Reinforcement Learning

In this paper the connection between stochastic optimal control and reinforcement learning is investigated. Our main motivation is to apply importance sampling to sampling rare events which can be reformulated as an optimal control problem. By using a parameterised approach the optimal control problem becomes a stochastic optimization problem which still raises some open questions regarding how to tackle the scalability to high-dimensional problems and how to deal with the intrinsic metastability of the system. To explore new methods we link the optimal control problem to reinforcement learning since both share the same underlying framework, namely a Markov Decision Process (MDP). For the optimal control problem we show how the MDP can be formulated. In addition we discuss how the stochastic optimal control problem can be interpreted in the framework of reinforcement learning. At the end of the article we present the application of two different reinforcement learning algorithms to the optimal control problem and a comparison of the advantages and disadvantages of the two algorithms.

math.OC

Improving control based importance sampling strategies for metastable diffusions via adapted metadynamics

Sampling rare events in metastable dynamical systems is often a computationally expensive task and one needs to resort to enhanced sampling methods such as importance sampling. Since we can formulate the problem of finding optimal importance sampling controls as a stochastic optimization problem, this then brings additional numerical challenges and the convergence of corresponding algorithms might as well suffer from metastabilty. In this article, we address this issue by combining systematic control approaches with the heuristic adaptive metadynamics method. Crucially, we approximate the importance sampling control by a neural network, which makes the algorithm in principle feasible for high-dimensional applications. We can numerically demonstrate in relevant metastable problems that our algorithm is more effective than previous attempts and that only the combination of the two approaches leads to a satisfying convergence and therefore to an efficient sampling in certain metastable settings.

math.OC

Learning Koopman eigenfunctions of stochastic diffusions with optimal importance sampling and ISOKANN

For stochastic diffusion processes the dominant eigenfunctions of the corresponding Koopman operator contain important information about the slow-scale dynamics, that is, about the location and frequency of rare events. In this article, we reformulate the eigenproblem in terms of $χ$-functions in the ISOKANN framework and discuss how optimal control and importance sampling allows for zero variance sampling of these functions. We provide a new formulation of the ISOKANN algorithm allowing for a proof of convergence and incorporate the optimal control result to obtain an adaptive iterative algorithm alternating between importance sampling and $χ$-function approximation. We demonstrate the usage of our proposed method in experiments increasing the approximation accuracy by several orders of magnitude.

math.DS

Extending Transition Path Theory: Periodically-Driven and Finite-Time Dynamics

Given two distinct subsets $A,B$ in the state space of some dynamical system, Transition Path Theory (TPT) was successfully used to describe the statistical behavior of transitions from $A$ to $B$ in the ergodic limit of the stationary system. We derive generalizations of TPT that remove the requirements of stationarity and of the ergodic limit, and provide this powerful tool for the analysis of other dynamical scenarios: periodically forced dynamics and time-dependent finite-time systems. This is partially motivated by studying applications such as climate, ocean, and social dynamics. On simple model examples we show how the new tools are able to deliver quantitative understanding about the statistical behavior of such systems. We also point out explicit cases where the more general dynamical regimes show different behaviors to their stationary counterparts, linking these tools directly to bifurcations in non-deterministic systems.

math.DS