SearcharxivSearch

arXiv subjects

Daniel Bartl

Publications and source records attributed to Daniel Bartl.

At least 19 recordsLinked to original sources

On the structure of marginals in high dimensions

Let $G, G_1,\dots,G_N$ be independent copies of a standard gaussian random vector in $\mathbb{R}^d$ and denote by $\Gamma = \sum_{i=1}^N \langle G_i,\cdot\rangle e_i$ the standard gaussian ensemble. We show that, for any set $A\subset S^{d-1}$, with exponentially high probability, \[ \sup_{x\in A} \frac{1}{N}\sum_{i=1}^N \big| (\Gamma x)^\sharp_i - q_i\big| \le c \frac{ \mathbb{E} \sup_{x\in A} \langle G,x\rangle + \log^2N }{\sqrt N }. \] Here each $q_i$ is the $\frac{i}{N+1}$-quantile of the standard normal distribution and $(\Gamma x)^\sharp $ denotes the monotone increasing rearrangement of the vector $\Gamma x$. The estimate is sharp up to a possible logarithmic factor and significantly extends previously known bounds. Moreover, we show that similar estimates hold in much greater generality: after replacing the gaussian quantiles by the appropriate ones, the same phenomenon persists for a broad class of random vectors.

math.PR

The geometry of the adapted Bures--Wasserstein space

The adapted Bures--Wasserstein space consists of Gaussian processes endowed with the adapted Wasserstein distance. It can be viewed as the analogue of the classical Bures--Wasserstein space in optimal transport for the setting of stochastic processes, where the standard Wasserstein distance is inadequate and has to be replaced by its adapted counterpart. We develop a comprehensive geometric theory for the adapted Bures--Wasserstein space, thereby also providing the first results on the fine geometric structure of adapted optimal transport. In particular, we show that the adapted Bures--Wasserstein space is an Alexandrov space with non-negative curvature and provide explicit descriptions of tangent cones and exponential maps. Moreover, we show that Gaussian processes satisfying a natural non-degeneracy condition form a geodesically convex subspace. This subspace is characterized precisely by the property that its tangent cones are linear and hence coincide with the tangent space.

math.PR

Fast Wasserstein rates for estimating probability distributions of probabilistic graphical models

Using i.i.d. data to estimate a high-dimensional distribution in Wasserstein distance is a fundamental instance of the curse of dimensionality. We explore how structural knowledge about the data-generating process which gives rise to the distribution can be used to overcome this curse. More precisely, we work with the set of distributions of probabilistic graphical models for a given directed acyclic graph. It turns out that this knowledge is only helpful if it can be quantified, which we formalize via smoothness conditions on the transition kernels in the disintegration corresponding to the graph. In this case, we prove that the rate of estimation is governed by the local structure of the graph, more precisely by dimensions corresponding to single nodes together with their parent nodes. The precise rate depends on the exact notion of smoothness assumed for the kernels, where either weak (Wasserstein-Lipschitz) or strong (bidirectional Total-Variation-Lipschitz) conditions lead to different results. We prove sharpness under the strong condition and show that this condition covers, as a special case, distributions having a positive Lipschitz density.

math.ST

Robust, sub-Gaussian mean estimators in metric spaces

Estimating the mean of a random vector from i.i.d. data has received considerable attention, and the optimal accuracy one may achieve with a given confidence is fairly well understood by now. When the data take values in more general metric spaces, an appropriate extension of the notion of the mean is the Fr\'echet mean. While asymptotic properties of the most natural Fr\'echet mean estimator (the empirical Fr\'echet mean) have been thoroughly researched, non-asymptotic performance bounds have only been studied recently. The aim of this paper is to study the performance of estimators of the Fr\'echet mean in general metric spaces under possibly heavy-tailed and contaminated data. In such cases, the empirical Fr\'echet mean is a poor estimator. We propose a general estimator based on high-dimensional extensions of trimmed means and prove general performance bounds. Unlike all previously established bounds, ours generalize the optimal bounds known for Euclidean data. The main message of the bounds is that, much like in the Euclidean case, the optimal accuracy is governed by two "variance" terms: a "global variance" term that is independent of the prescribed confidence, and a potentially much smaller, confidence-dependent "local variance" term. We apply our results for metric spaces with curvature bounded from below, such as Wasserstein spaces, and for uniformly convex Banach spaces.

math.ST

Uniform mean estimation via generic chaining

We introduce an empirical functional $\Psi$ that is an optimal uniform mean estimator: Let $F\subset L_2(\mu)$ be a class of mean zero functions, $u$ is a real valued function, and $X_1,\dots,X_N$ are independent, distributed according to $\mu$. We show that under minimal assumptions, with $\mu^{\otimes N}$ exponentially high probability, \[ \sup_{f\in F} |\Psi(X_1,\dots,X_N,f) - \mathbb{E} u(f(X))| \leq c R(F) \frac{ \mathbb{E} \sup_{f\in F } |G_f| }{\sqrt N}, \] where $(G_f)_{f\in F}$ is the gaussian processes indexed by $F$ and $R(F)$ is an appropriate notion of `diameter' of the class $\{u(f(X)) : f\in F\}$. The fact that such a bound is possible is surprising, and it leads to the solution of various key problems in high dimensional probability and high dimensional statistics. The construction is based on combining Talagrand's generic chaining mechanism with optimal mean estimation procedures for a single real-valued random variable.

math.PR

Do we really need the Rademacher complexities?

We study the fundamental problem of learning with respect to the squared loss in a convex class. The state-of-the-art sample complexity estimates in this setting rely on Rademacher complexities, which are generally difficult to control. We prove that, contrary to prevailing belief and under minimal assumptions, the sample complexity is not governed by the Rademacher complexities but rather by the behaviour of the limiting gaussian process. In particular, all such learning problems that have the same $L_2$-structure -- even those with heavy-tailed distributions -- share the same sample complexity. This constitutes the first universality result for general convex learning problems. The proof is based on a novel learning procedure, and its performance is studied by combining optimal mean estimation techniques for real-valued random variables with Talagrand's generic chaining method.

math.ST

The Wasserstein Space of Stochastic Processes in Continuous Time

Researchers from different areas have independently defined extensions of the usual weak convergence of laws of stochastic processes with the goal of adequately accounting for the flow of information. Natural approaches are convergence of the Aldous--Knight prediction process, Hellwig's information topology, convergence in adapted distribution in the sense of Hoover--Keisler and the weak topology induced by optimal stopping problems. The first main contribution of this article is that on continuous processes with natural filtrations there exists a canonical adapted weak topology which can be defined by all of these approaches; moreover, the adapted weak topology is metrized by a suitable adapted Wasserstein distance $\mathcal{AW}$. While the set of processes with natural filtrations is not complete, we establish that its completion consists precisely of the space ${\rm FP}$ of stochastic processes with general filtrations. We also show that $({\rm FP}, \mathcal{AW})$ exhibits several desirable properties. Specifically, it is Polish, martingales form a closed subset and approximation results such as Donsker's theorem extend to $\mathcal{AW}$.

math.PR

Optimal nonparametric estimation of the expected shortfall risk

We address the problem of estimating the expected shortfall risk of a financial loss using a finite number of i.i.d. data. It is well known that the classical plug-in estimator suffers from poor statistical performance when faced with (heavy-tailed) distributions that are commonly used in financial contexts. Further, it lacks robustness, as the modification of even a single data point can cause a significant distortion. We propose a novel procedure for the estimation of the expected shortfall and prove that it recovers the best possible statistical properties (dictated by the central limit theorem) under minimal assumptions and for all finite numbers of data. Further, this estimator is adversarially robust: even if a (small) proportion of the data is maliciously modified, the procedure continuous to optimally estimate the true expected shortfall risk. We demonstrate that our estimator outperforms the classical plug-in estimator through a variety of numerical experiments across a range of standard loss distributions.

q-fin.RM

Numerical method for nonlinear Kolmogorov PDEs via sensitivity analysis

We examine nonlinear Kolmogorov partial differential equations (PDEs). Here the nonlinear part of the PDE comes from its Hamiltonian where one maximizes over all possible drift and diffusion coefficients which fall within a $\varepsilon$-neighborhood of pre-specified baseline coefficients. Our goal is to quantify and compute how sensitive those PDEs are to such a small nonlinearity, and then use the results to develop an efficient numerical method for their approximation. We show that as $\varepsilon\downarrow 0$, the nonlinear Kolmogorov PDE equals the linear Kolmogorov PDE defined with respect to the corresponding baseline coefficients plus $\varepsilon$ times a correction term which can be also characterized by the solution of another linear Kolmogorov PDE involving the baseline coefficients. As these linear Kolmogorov PDEs can be efficiently solved in high-dimensions by exploiting their Feynman-Kac representation, our derived sensitivity analysis then provides a Monte Carlo based numerical method which can efficiently solve these nonlinear Kolmogorov equations. We establish an error and complexity analysis for our numerical method. Moreover, we provide numerical examples in up to 100 dimensions to empirically demonstrate the applicability of our numerical method.

math.NA

A uniform Dvoretzky-Kiefer-Wolfowitz inequality

We show that under minimal assumptions on a class of functions $\mathcal{H}$ defined on a probability space $(\mathcal{X},\mu)$, there is a threshold $\Delta_0$ satisfying the following: for every $\Delta\geq\Delta_0$, with probability at least $1-2\exp(-c\Delta m)$ with respect to $\mu^{\otimes m}$, \[ \sup_{h\in\mathcal{H}} \sup_{t\in\mathbb{R}} \left| \mathbb{P}(h(X)\leq t) - \frac{1}{m}\sum_{i=1}^m 1_{(-\infty,t]}(h(X_i)) \right| \leq \sqrt{\Delta};\] here $X$ is distributed according to $\mu$ and $(X_i)_{i=1}^m$ are independent copies of $X$. The value of $\Delta_0$ is determined by an unexpected complexity parameter of the class $\mathcal{H}$ that captures the set's geometry (Talagrand's $\gamma_1$-functional). The bound, the probability estimate and the value of $\Delta_0$ are all optimal up to a logarithmic factor.

math.PR

Sensitivity of robust optimization problems under drift and volatility uncertainty

We examine optimization problems in which an investor has the opportunity to trade in $d$ stocks with the goal of maximizing her worst-case cost of cumulative gains and losses. Here, worst-case refers to taking into account all possible drift and volatility processes for the stocks that fall within a $\varepsilon$-neighborhood of predefined fixed baseline processes. Although solving the worst-case problem for a fixed $\varepsilon>0$ is known to be very challenging in general, we show that it can be approximated as $\varepsilon\to 0$ by the baseline problem (computed using the baseline processes) in the following sense: Firstly, the value of the worst-case problem is equal to the value of the baseline problem plus $\varepsilon$ times a correction term. This correction term can be computed explicitly and quantifies how sensitive a given optimization problem is to model uncertainty. Moreover, approximately optimal trading strategies for the worst-case problem can be obtained using optimal strategies from the corresponding baseline problem.

math.OC

Optimal non-gaussian Dvoretzky-Milman embeddings

We construct the first non-gaussian ensemble that yields the optimal estimate in the Dvoretzky-Milman Theorem: the ensemble exhibits almost Euclidean sections in arbitrary normed spaces of the same dimension as the gaussian embedding -- despite being very far from gaussian (in fact, it happens to be heavy-tailed).

math.FA

Empirical approximation of the gaussian distribution in $\mathbb{R}^d$

Let $G_1,\dots,G_m$ be independent copies of the standard gaussian random vector in $\mathbb{R}^d$. We show that there is an absolute constant $c$ such that for any $A \subset S^{d-1}$, with probability at least $1-2\exp(-c\Delta m)$, for every $t\in\mathbb{R}$, \[ \sup_{x \in A} \left| \frac{1}{m}\sum_{i=1}^m 1_{ \{\langle G_i,x\rangle \leq t \}} - \mathbb{P}(\langle G,x\rangle \leq t) \right| \leq \Delta + \sigma(t) \sqrt\Delta. \] Here $\sigma(t) $ is the variance of $1_{\{\langle G,x\rangle\leq t\}}$ and $\Delta\geq \Delta_0$, where $\Delta_0$ is determined by an unexpected complexity parameter of $A$ that captures the set's geometry (Talagrand's $\gamma_1$ functional). The bound, the probability estimate, and the value of $\Delta_0$ are all (almost) optimal. We use this fact to show that if $\Gamma=\sum_{i=1}^m \langle G_i,x\rangle e_i$ is the random matrix that has $G_1,\dots,G_m$ as its rows, then the structure of $\Gamma(A)=\{\Gamma x: x\in A\}$ is far more rigid and well-prescribed than was previously expected.

math.PR

The Wasserstein space of stochastic processes

Wasserstein distance induces a natural Riemannian structure for the probabilities on the Euclidean space. This insight of classical transport theory is fundamental for tremendous applications in various fields of pure and applied mathematics. We believe that an appropriate probabilistic variant, the adapted Wasserstein distance AW, can play a similar role for the class FP of filtered processes, i.e. stochastic processes together with a filtration. In contrast to other topologies for stochastic processes, probabilistic operations such as the Doob-decomposition, optimal stopping and stochastic control are continuous w.r.t. AW. We also show that (FP,AW) is a geodesic space, isometric to a classical Wasserstein space, and that martingales form a closed geodesically convex subspace.

math.PR

On a variance dependent Dvoretzky-Kiefer-Wolfowitz inequality

Let $X$ be a real-valued random variable with distribution function $F$. Set $X_1,\dots, X_m$ to be independent copies of $X$ and let $F_m$ be the corresponding empirical distribution function. We show that there are absolute constants $c_0$ and $c_1$ such that if $Δ\geq c_0\frac{\log\log m}{m}$, then with probability at least $1-2\exp(-c_1Δm)$, for every $t\in\mathbb{R}$ that satisfies $F(t)\in[Δ,1-Δ]$, \[ |F_m(t) - F(t) | \leq \sqrt{Δ\min\{F(t),1-F(t)\} } .\] Moreover, this estimate is optimal up to the multiplicative constants $c_0$ and $c_1$.

math.PR

Sensitivity of multiperiod optimization problems in adapted Wasserstein distance

We analyze the effect of small changes in the underlying probabilistic model on the value of multi-period stochastic optimization problems and optimal stopping problems. We work in finite discrete time and measure these changes with the adapted Wasserstein distance. We prove explicit first-order approximations for both problems. Expected utility maximization is discussed as a special case.

math.OC

Non-asymptotic convergence rates for the plug-in estimation of risk measures

Let $ρ$ be a general law--invariant convex risk measure, for instance the average value at risk, and let $X$ be a financial loss, that is, a real random variable. In practice, either the true distribution $μ$ of $X$ is unknown, or the numerical computation of $ρ(μ)$ is not possible. In both cases, either relying on historical data or using a Monte-Carlo approach, one can resort to an i.i.d.\ sample of $μ$ to approximate $ρ(μ)$ by the finite sample estimator $ρ(μ_N)$ (where $μ_N$ denotes the empirical measure of $μ$). In this article we investigate convergence rates of $ρ(μ_N)$ to $ρ(μ)$. We provide non-asymptotic convergence rates for both the deviation probability and the expectation of the estimation error. The sharpness of these convergence rates is analyzed. Our framework further allows for hedging, and the convergence rates we obtain depend neither on the dimension of the underlying assets, nor on the number of options available for trading.

q-fin.RM

Structure preservation via the Wasserstein distance

We show that under minimal assumptions on a random vector $X\in\mathbb{R}^d$ and with high probability, given $m$ independent copies of $X$, the coordinate distribution of each vector $(\langle X_i,\theta \rangle)_{i=1}^m$ is dictated by the distribution of the true marginal $\langle X,\theta \rangle$. Specifically, we show that with high probability, \[\sup_{\theta \in S^{d-1}} \left( \frac{1}{m}\sum_{i=1}^m \left|\langle X_i,\theta \rangle^\sharp - \lambda^\theta_i \right|^2 \right)^{1/2} \leq c \left( \frac{d}{m} \right)^{1/4},\] where $\lambda^{\theta}_i = m\int_{(\frac{i-1}{m}, \frac{i}{m}]} F_{ \langle X,\theta \rangle }^{-1}(u)\,du$ and $a^\sharp$ denotes the monotone non-decreasing rearrangement of $a$. Moreover, this estimate is optimal. The proof follows from a sharp estimate on the worst Wasserstein distance between a marginal of $X$ and its empirical counterpart, $\frac{1}{m} \sum_{i=1}^m \delta_{\langle X_i, \theta \rangle}$.

math.ST