SearcharxivSearch

arXiv subjects

Johannes Wiesel

Publications and source records attributed to Johannes Wiesel.

At least 19 recordsLinked to original sources

Bounding adapted Wasserstein metrics

The Wasserstein distance $\mathcal{W}_p$ is an important instance of an optimal transport cost. Its numerous mathematical properties as well as applications to various fields such as mathematical finance and statistics have been well studied in recent years. The adapted Wasserstein distance $\mathcal{A}\mathcal{W}_p$ extends this theory to laws of discrete time stochastic processes in their natural filtrations, making it particularly well suited for analyzing time-dependent stochastic optimization problems. While the topological differences between $\mathcal{A}\mathcal{W}_p$ and $\mathcal{W}_p$ are well understood, their differences as metrics remain largely unexplored beyond the trivial bound $\mathcal{W}_p\lesssim \mathcal{A}\mathcal{W}_p$. This paper closes this gap by providing upper bounds of $\mathcal{A}\mathcal{W}_p$ in terms of $\mathcal{W}_p$ through investigation of the smooth adapted Wasserstein distance. Our upper bounds are explicit and are given by a sum of $\mathcal{W}_p$, Eder's modulus of continuity and a term characterizing the tail behavior of measures. As a consequence, upper bounds on $\mathcal{W}_p$ automatically hold for $\mathcal{AW}_p$ under mild regularity assumptions on the measures considered. A particular instance of our findings is the inequality $\mathcal{A}\mathcal{W}_1\le C\sqrt{\mathcal{W}_1}$ on the set of measures that have Lipschitz kernels. Our work also reveals how smoothing of measures affects the adapted weak topology. In fact, we find that the topology induced by the smooth adapted Wasserstein distance exhibits a non-trivial interpolation property, which we characterize explicitly: it lies in between the adapted weak topology and the weak topology, and the inclusion is governed by the decay of the smoothing parameter.

math.PR

Almost stochastic dominance via optimal transport

We study parametric classes of almost stochastic dominance on general Polish spaces as order relations for probability distributions with a parameter $γ\in [0,1]$. Larger values of $γ$ correspond to weaker order relations: $γ=0$ gives classical stochastic dominance $\le_{st}$, whereas $γ=1$ gives a complete preorder based on comparison of expectations of a fixed increasing function $g$. It is well known that $X \le_{st} Y$ can be characterized by the existence of a solution to an optimal transport problem with $\mathrm{OT}_c(X,Y)=0$ for a suitable cost function $c$. We generalize this idea so that the best possible parameter $γ$ for almost stochastic dominance can be determined from the solution of an optimal transport problem. Using a generalization of the classical Kantorovich--Rubinstein duality theorem to quasi-pseudo-metrics, we derive a dual characterization of the order in terms of expectation comparisons for a parametric class of test functions. Consequently, our relations are always transitive, in contrast to some other recent approaches to almost stochastic dominance based on optimal transport. A natural multivariate approach to almost stochastic dominance, based on classes of test functions with bounds on partial derivatives, was recently introduced by Müller et al. (2025). We show that this approach is a special case of our framework and derive the best possible parameters $γ$ for examples considered there, as well as for other examples from the literature. We also prove a robustness result showing that, under small perturbations of the distributions in a Wasserstein-type metric related to the optimal transport problem, the best possible $γ$ increases only slightly.

math.PR

Adapted Wasserstein Barycenters of Gaussian Processes

We study barycenters of filtered Gaussian processes in adapted Wasserstein space. The adapted Wasserstein distance refines classical optimal transport by requiring transport plans to respect the temporal flow of information, making it the natural metric for stochastic systems with filtration constraints, as in stochastic control, mathematical finance, and sequential decision problems. We prove that the \emph{unrestricted} barycenter problem for weighted Fréchet means of filtered Gaussian inputs admits a solution with Gaussian underlying law, representable as an enlarged filtered Gaussian process but not necessarily as an ordinary one. The problem decomposes into finitely many classical Bures--Wasserstein barycenter problems for the covariance contributions of the successive innovations. We then treat the \emph{restricted} problem, in which the barycenter is required to be an ordinary filtered Gaussian process, giving a rank and common-noise criterion for when the two problems agree, sufficient conditions for uniqueness, and first order optimality and regularity results. Under a martingale constraint we obtain an explicit solution via martingale projection and Bures--Wasserstein barycenters of the Gaussian increments. Beyond their intrinsic theoretical interest, our results provide a principled way to build representative models from collections of Gaussian stochastic systems, with applications to stochastic optimization, robust finance, and sequential statistical analysis.

math.PR

Dependence Measures via Adapted Optimal Transport: Stability and Rates of Convergence

Recently studied dependence measures, such as Chatterjee's rank correlation, that characterize both independence and perfect functional dependence, provide a powerful framework for detecting nonlinear dependencies. However, these measures cannot be weakly continuous, which limits the applicability of classical plug-in estimators based on empirical distributions. This obstruction is natural, as such measures are defined via conditional distributions and not through their joint law alone. In this paper, we introduce an optimal transport-based mode of convergence that captures weak convergence of conditional distributions and restores continuity for a broad class of dependence measures. We relate this mode of convergence to the adapted Wasserstein distance, the Knothe-Rosenblatt distance and the d1-metric on copulas. Building on this perspective, we propose a copula estimator based on the adapted empirical measure and compare it with the classical rank-based checkerboard estimator. For both estimators, we derive O(N^{-1/3})-rates of convergence with respect to metrics that capture conditional weak continuity. As a consequence, we obtain the same rates for plug-in estimators of several classes of dependence measures, including rank-based and rearranged dependence measures.

math.ST

The fast rate of convergence of the smooth adapted Wasserstein distance

Estimating a $d$-dimensional distribution $μ$ by the empirical measure $\hatμ_n$ of its samples is an important task in probability theory, statistics and machine learning. It is well known that $\mathbb{E}[\mathcal{W}_p(\hatμ_n, μ)]\lesssim n^{-1/d}$ for $d>2p$, where $\mathcal{W}_p$ denotes the $p$-Wasserstein metric. An effective tool to combat this curse of dimensionality is the smooth Wasserstein distance $\mathcal{W}^{(σ)}_p$, which measures the distance between two probability measures after having convolved them with isotropic Gaussian noise $\mathcal{N}(0,σ^2\text{I})$. In this paper we apply this smoothing technique to the adapted Wasserstein distance. We show that the smooth adapted Wasserstein distance $\mathcal{A}\mathcal{W}_p^{(σ)}$ achieves the fast rate of convergence $\mathbb{E}[\mathcal{A}\mathcal{W}_p^{(σ)}(\hatμ_n, μ)]\lesssim n^{-1/2}$, if $μ$ is subgaussian. This result follows from the surprising fact, that any subgaussian measure $μ$ convolved with a Gaussian distribution has locally Lipschitz kernels.

math.PR

Sample complexity for divergence regularized optimal transport with radial cost

We prove a new sample complexity result for divergence regularized optimal transport. Our bound holds for probability measures on~$\mathbb{R}^d$ with exponential tail decay and for radial cost functions that satisfy a local Lipschitz condition. It is sharp up to logarithmic factors, and captures the intrinsic dimension of the marginal distributions through a generalized covering number of their supports. Examples that fit into our framework include subexponential and subgaussian distributions and radial cost functions $c(x,y)=|x-y|^p$ for $p\ge 1$ with logarithmic entropy or polynomial $α$-divergence.

math.ST

Convergence of the adapted empirical measure for mixing observations

The adapted Wasserstein distance $\mathcal{AW}$ is a modification of the classical Wasserstein metric, that provides robust and dynamically consistent comparisons of laws of stochastic processes, and has proved particularly useful in the analysis of stochastic control problems, model uncertainty, and mathematical finance. In applications, the law of a stochastic process $μ$ is not directly observed, and has to be inferred from a finite number of samples. As the empirical measure is not $\mathcal{AW}$-consistent, Backhoff, Bartl, Beiglböck and Wiesel introduced the adapted empirical measure $\widehatμ^N$, a suitable modification, and proved its $\mathcal{AW}$-consistency when observations are i.i.d. In this paper we study $\mathcal{AW}$-convergence of the adapted empirical measure $\widehatμ^N$ to the population distribution $μ$, for observations satisfying a generalization of the $η$-mixing condition introduced by Kontorovich and Ramanan. We establish moment bounds and sub-exponential concentration inequalities for $\mathcal{AW}(μ,\widehatμ^N)$, and prove consistency of $\widehatμ^N$. In addition, we extend the Bounded Differences inequality of Kontorovich and Ramanan for $η$-mixing observations to uncountable spaces, a result that may be of independent interest. Numerical simulations illustrating our theory are also provided.

math.PR

Empirical martingale projections via the adapted Wasserstein distance

Given a collection of multidimensional pairs $\{(X_i,Y_i):1 \leq i\leq n\}$, we study the problem of projecting the associated suitably smoothed empirical measure onto the space of martingale couplings (i.e. distributions satisfying $\mathbb{E}[Y|X]=X$) using the adapted Wasserstein distance. We call the resulting distance the smoothed empirical martingale projection distance (SE-MPD), for which we obtain an explicit characterization. We also show that the space of martingale couplings remains invariant under the smoothing operation. We study the asymptotic limit of the SE-MPD, which converges at a parametric rate as the sample size increases if the pairs are either i.i.d. or satisfy appropriate mixing assumptions. Additional finite-sample results are also investigated. Using these results, we introduce a novel consistent martingale coupling hypothesis test, which we apply to test the existence of arbitrage opportunities in recently introduced neural network-based generative models for asset pricing calibration.

math.PR

On the Martingale Schrödinger Bridge between Two Distributions

We study a martingale Schrödinger bridge problem: given two probability distributions, find their martingale coupling with minimal relative entropy. Our main result provides Schrödinger potentials for this coupling. Namely, under certain conditions, the log-density of the optimal coupling is given by a triplet of real functions representing the marginal and martingale constraints. The potentials are also described as the solution of a dual problem.

math.PR

A dynamic programming principle for multiperiod control problems with bicausal constraints

We consider multiperiod stochastic control problems with non-parametric uncertainty on the underlying probabilistic model. We derive a new metric on the space of probability measures, called the adapted $(p, \infty)$--Wasserstein distance $\mathcal{AW}_p^\infty$ with the following properties: (1) the adapted $(p, \infty)$--Wasserstein distance generates a topology that guarantees continuity of stochastic control problems and (2) the corresponding $\mathcal{AW}_p^\infty$-distributionally robust optimization (DRO) problem can be computed via a dynamic programming principle involving one-step Wasserstein-DRO problems. If the cost function is semi-separable, then we further show that a minimax theorem holds, even though balls with respect to $\mathcal{AW}_p^\infty$ are neither convex nor compact in general. We also derive first-order sensitivity results.

math.OC

Sparsity of Quadratically Regularized Optimal Transport: Bounds on concentration and bias

We study the quadratically regularized optimal transport (QOT) problem for quadratic cost and compactly supported marginals $μ$ and $ν$. It has been empirically observed that the optimal coupling $π_ε$ for the QOT problem has sparse support for small regularization parameter $ε>0.$ In this article we provide the first quantitative description of this phenomenon in general dimension: we derive bounds on the size and on the location of the support of $π_ε$ compared to the Monge coupling. Our analysis is based on pointwise bounds on the density of $π_ε$ together with Minty's trick, which provides a quadratic detachment from the optimal transport duality gap. In the self-transport setting $μ=ν$ we obtain optimal rates of order $ε^{\frac{1}{2+d}}.$

math.OC

Max-sliced Wasserstein concentration and uniform ratio bounds of empirical measures on RKHS

Optimal transport and the Wasserstein distance $\mathcal{W}_p$ have recently seen a number of applications in the fields of statistics, machine learning, data science, and the physical sciences. These applications are however severely restricted by the curse of dimensionality, meaning that the number of data points needed to estimate these problems accurately increases exponentially in the dimension. To alleviate this problem, a number of variants of $\mathcal{W}_p$ have been introduced. We focus here on one of these variants, namely the max-sliced Wasserstein metric $\overline{\mathcal{W}}_p$. This metric reduces the high-dimensional minimization problem given by $\mathcal{W}_p$ to a maximum of one-dimensional measurements in an effort to overcome the curse of dimensionality. In this note we derive concentration results and upper bounds on the expectation of $\overline{\mathcal{W}}_p$ between the true and empirical measure on unbounded reproducing kernel Hilbert spaces. We show that, under quite generic assumptions, probability measures concentrate uniformly fast in one-dimensional subspaces, at (nearly) parametric rates. Our results rely on an improvement of currently known bounds for $\overline{\mathcal{W}}_p$ in the finite-dimensional case.

math.ST

The out-of-sample prediction error of the square-root-LASSO and related estimators

We study the classical problem of predicting an outcome variable, $Y$, using a linear combination of a $d$-dimensional covariate vector, $\mathbf{X}$. We are interested in linear predictors whose coefficients solve: % \begin{align*} \inf_{\boldsymbolβ \in \mathbb{R}^d} \left( \mathbb{E}_{\mathbb{P}_n} \left[ \left(Y-\mathbf{X}^{\top}β\right)^r \right] \right)^{1/r} +δ\, ρ\left(\boldsymbolβ\right), \end{align*} where $δ>0$ is a regularization parameter, $ρ:\mathbb{R}^d\to \mathbb{R}_+$ is a convex penalty function, $\mathbb{P}_n$ is the empirical distribution of the data, and $r\geq 1$. We present three sets of new results. First, we provide conditions under which linear predictors based on these estimators % solve a \emph{distributionally robust optimization} problem: they minimize the worst-case prediction error over distributions that are close to each other in a type of \emph{max-sliced Wasserstein metric}. Second, we provide a detailed finite-sample and asymptotic analysis of the statistical properties of the balls of distributions over which the worst-case prediction error is analyzed. Third, we use the distributionally robust optimality and our statistical analysis to present i) an oracle recommendation for the choice of regularization parameter, $δ$, that guarantees good out-of-sample prediction error; and ii) a test-statistic to rank the out-of-sample performance of two different linear estimators. None of our results rely on sparsity assumptions about the true data generating process; thus, they broaden the scope of use of the square-root lasso and related estimators in prediction problems.

math.ST

On concentration of the empirical measure for radial transport costs

Let $μ$ be a probability measure on $\mathbb{R}^d$ and $μ_N$ its empirical measure with sample size $N$. We prove a concentration inequality for the optimal transport cost between $μ$ and $μ_N$ for radial cost functions with polynomial local growth, that can have superpolynomial global growth. This result generalizes and improves upon estimates of Fournier and Guillin. The proof combines ideas from empirical process theory with known concentration rates for compactly supported $μ$. By partitioning $\mathbb{R}^d$ into annuli, we infer a global estimate from local estimates on the annuli and conclude that the global estimate can be expressed as a sum of the local estimate and a mean-deviation probability for which efficient bounds are known.

math.ST

Strategies with minimal norm are optimal for expected utility maximization under high model ambiguity

We investigate an expected utility maximization problem under model uncertainty in a one-period financial market. We capture model uncertainty by replacing the baseline model $\mathbb{P}$ with an adverse choice from a Wasserstein ball of radius $k$ around $\mathbb{P}$ in the space of probability measures and consider the corresponding Wasserstein distributionally robust optimization problem. We show that optimal solutions converge to a strategy with minimal norm when uncertainty is increasingly large, i.e. when the radius $k$ tends to infinity.

math.OC

Sensitivity of multiperiod optimization problems in adapted Wasserstein distance

We analyze the effect of small changes in the underlying probabilistic model on the value of multi-period stochastic optimization problems and optimal stopping problems. We work in finite discrete time and measure these changes with the adapted Wasserstein distance. We prove explicit first-order approximations for both problems. Expected utility maximization is discussed as a special case.

math.OC

An optimal transport based characterization of convex order

For probability measures $μ,ν$ and $ρ$ define the cost functionals \begin{align*} C(μ,ρ):=\sup_{π\in Π(μ,ρ)} \int \langle x,y\rangle\, π(dx,dy),\quad C(ν,ρ):=\sup_{π\in Π(ν,ρ)} \int \langle x,y\rangle\, π(dx,dy), \end{align*} where $\langle\cdot, \cdot\rangle$ denotes the scalar product and $Π(\cdot,\cdot)$ is the set of couplings. We show that two probability measures $μ$ and $ν$ on $\mathbb{R}^d$ with finite first moments are in convex order (i.e. $μ\preceq_cν$) iff $C(μ,ρ)\le C(ν,ρ)$ holds for all probability measures $ρ$ on $\mathbb{R}^d$ with bounded support. This generalizes a result by Carlier. Our proof relies on a quantitative bound for the infimum of $\int f\,dν-\int f\,dμ$ over all $1$-Lipschitz functions $f$, which is obtained through optimal transport duality and Brenier's theorem. Building on this result, we derive new proofs of well-known one-dimensional characterizations of convex order. We also describe new computational methods for investigating convex order and applications to model-independent arbitrage strategies in mathematical finance.

math.PR