SearcharxivSearch

arXiv subjects

Julian Gerstenberg

Publications and source records attributed to Julian Gerstenberg.

6 recordsLinked to original sources

On Policy Evaluation Algorithms in Distributional Reinforcement Learning

We introduce a novel class of algorithms to efficiently approximate the unknown return distributions in policy evaluation problems from distributional reinforcement learning (DRL). The proposed distributional dynamic programming algorithms are suitable for underlying Markov decision processes (MDPs) having an arbitrary probabilistic reward mechanism, including continuous reward distributions with unbounded support being potentially heavy-tailed. For a plain instance of our proposed class of algorithms we prove error bounds, both within Wasserstein and Kolmogorov--Smirnov distances. Furthermore, for return distributions having probability density functions the algorithms yield approximations for these densities; error bounds are given within supremum norm. We introduce the concept of quantile-spline discretizations to come up with algorithms showing promising results in simulation experiments. While the performance of our algorithms can rigorously be analysed they can be seen as universal black box algorithms applicable to a large class of MDPs. We also derive new properties of probability metrics commonly used in DRL on which our quantitative analysis is based.

stat.ML

Exchangeable Laws in Borel Data Structures

Motivated by statistical practice, category theory terminology is used to introduce Borel data structures and study exchangeability in an abstract framework. A generalization of de Finetti's theorem is shown and natural transformations are used to present functional representation theorems (FRTs). Proofs of the latter are based on a classical result by D.N.Hoover providing a functional representation for exchangeable arrays indexed by finite tuples of integers, together with an universality result for Borel data structures. A special class of Borel data structures are array-type data structures, which are introduced using the novel concept of an indexing system. Studying natural transformations mapping into arrays gives explicit versions of FRTs, which in examples coincide with well-known Aldous-Hoover-Kallenberg-type FRTs for (jointly) exchangeable arrays. The abstract "index arithmetic" presented unifies and generalizes technical arguments commonly encountered in the literature on exchangeability theory. Finally, the category theory approach is used to outline how an abstract notion of seperate exchangeability can be derived, again motivated from statistical practice.

math.PR

On solutions of the distributional Bellman equation

In distributional reinforcement learning not only expected returns but the complete return distributions of a policy are taken into account. The return distribution for a fixed policy is given as the solution of an associated distributional Bellman equation. In this note we consider general distributional Bellman equations and study existence and uniqueness of their solutions as well as tail properties of return distributions. We give necessary and sufficient conditions for existence and uniqueness of return distributions and identify cases of regular variation. We link distributional Bellman equations to multivariate affine distributional equations. We show that any solution of a distributional Bellman equation can be obtained as the vector of marginal laws of a solution to a multivariate affine distributional equation. This makes the general theory of such equations applicable to the distributional reinforcement learning setting.

stat.ML

General Erased-Word Processes: Product-Type Filtrations, Ergodic Laws and Martin Boundaries

We study the dynamics of erasing randomly chosen letters from words by introducing a certain class of discrete-time stochastic processes, general erased-word processes(GEWPs), and investigating three closely related topics: Representation, Martin boundary and filtration theory. We use de Finetti's theorem and the random exchangeable linear order to obtain a de Finetti-type representation of GEWPs involving induced order statistics. Our studies expose connections between exchangeability theory and certain poly-adic filtrations that can be found in other exchangeable random objects as well. We show that ergodic GEWPs generate backward filtrations of product-type and by that generalize a result by S.Laurent.

math.PR

Exchangeable interval hypergraphs and limits of ordered discrete structures

A hypergraph $(V,E)$ is called an interval hypergraph if there exists a linear order $l$ on $V$ such that every edge $e\in E$ is an interval w.r.t. $l$; we also assume that $\{j\}\in E$ for every $j\in V$. Our main result is a de Finetti-type representation of random exchangeable interval hypergraphs on $\mathbb{N}$ (EIHs): the law of every EIH can be obtained by sampling from some random compact subset $K$ of the triangle $\{(x,y):0\leq x\leq y\leq 1\}$ at iid uniform positions $U_1,U_2,\dots$, in the sense that, restricted to the node set $[n]:=\{1,\dots,n\}$ every non-singleton edge is of the form $e=\{i\in[n]:x<U_i<y\}$ for some $(x,y)\in K$. We obtain this result via the study of a related class of stochastic objects: erased-interval processes (EIPs). These are certain transient Markov chains $(I_n,η_n)_{n\in\mathbb{N}}$ such that $I_n$ is an interval hypergraph on $V=[n]$ w.r.t. the usual linear order (called interval system). We present an almost sure representation result for EIPs. Attached to each transient Markov chain is the notion of Martin boundary. The points in the boundary attached to EIPs can be seen as limits of growing interval systems. We obtain a one-to-one correspondence between these limits and compact subsets $K$ of the triangle with $(x,x)\in K$ for all $x\in[0,1]$. Interval hypergraphs are a generalizations of hierarchies and as a consequence we obtain a representation result for exchangeable hierarchies, which is close to a result of Forman, Haulk and Pitman. Several ordered discrete structures can be seen as interval systems with additional properties, i.e. Schröder trees and binary trees. We describe limits of Schröder trees as certain tree-like compact sets. Considering binary trees we thus obtain a homeomorphic description of the Martin boundary of Rémy's tree growth chain, which has been analyzed by Evans, Grübel and Wakolbinger.

math.PR

A boundary theory approach to de Finetti's theorem

We show that boundary theory for transient Markov chains, as initiated by Doob, can be used to prove de Finetti's classical representation result for exchangeable random sequences. We also include the relevant parts of the theory, with full proofs.

math.PR