SearcharxivSearch

arXiv subjects

Bernhard Schmitzer

Publications and source records attributed to Bernhard Schmitzer.

At least 19 recordsLinked to original sources

Monte Carlo Event Generation with Continuous Normalizing Flows

We apply Continuous Normalizing Flows trained with the Flow Matching method to the problem of phase-space sampling in Monte Carlo event generation for high-energy collider physics. Focusing on lepton-pair and top quark pair production with multiple jets, the two computationally most expensive processes at the Large Hadron Collider, we train helicity-conditioned Continuous Normalizing Flows to remap the random numbers used in matrix element evaluation. Compared to standard methods, we achieve unweighting efficiency improvements by factors of up to 184 and 25 for the two processes at their respective highest jet number, at the cost of an increased evaluation time. When combining the advantages of Continuous Normalizing Flows with the fast evaluation times of Coupling Layer based Flows, using the RegFlow approach, we find parton-level unweighted event generation walltime gains of about a factor of ten at the highest jet numbers. These substantial gains highlight the promise of samplers based on machine learning for next-generation collider experiments.

hep-ph

Density estimation from batched broken random samples

The broken random sample problem was first introduced by DeGroot, Feder, and Gole (1971, Ann. Math. Statist.): in each observation (batch), a random sample of $M$ i.i.d. point pairs $ ((X_i,Y_i))_{i=1}^M$ is drawn from a joint distribution with density $p(x,y)$, but we can observe only the unordered multisets $(X_i)_{i=1}^M$ and $(Y_i)_{i=1}^M$ separately; that is, the pairing information is lost. For large $M$, inferring $p$ from a single observation has been shown to be essentially impossible. In this paper, we propose a parametric method based on a pseudo-log-likelihood to estimate $p$ from $N$ i.i.d. broken sample batches, and we prove a fast convergence rate in $N$ for our estimator that is uniform in $M$, under mild assumptions.

math.ST

Efficient many-jet event generation with Flow Matching

We apply for the first time the Flow Matching method to the problem of phase-space sampling for event generation in high-energy collider physics. By training the model to remap the random numbers used to generate the momenta and helicities of the scattering matrix elements as implemented in the portable partonic event generator Pepper, we find substantial efficiency improvements in the studied processes. We focus our study on the highest final-state multiplicities in Drell--Yan and top--antitop pair production used in simulated samples for the Large Hadron Collider, which computationally are the most relevant ones. We find that the unweighting efficiencies improve by factors of 150 and 17, respectively, when compared to the standard approach of using a Vegas-based optimisation. We also compare Continuous Normalizing Flows trained with Flow Matching against the previously studied Normalizing Flows based on Coupling Layers and find that the former leads to better results, faster training and a better scaling behaviour across the studied multiplicity range.

hep-ph

Entropic transfer operators for stochastic systems

Dynamical systems can be analyzed via their Frobenius-Perron transfer operator and its estimation from data is an active field of research. Recently entropic transfer operators have been introduced to estimate the operator of deterministic systems. The approach is based on the regularizing properties of entropic optimal transport plans. In this article we generalize the method to stochastic and non-stationary systems and give a quantitative convergence analysis of the empirical operator as the available samples increase. We introduce a way to extend the operator's eigenfunctions to previously unseen samples, such that they can be efficiently included into a spectral embedding. The practicality and numerical scalability of the method are demonstrated on a real-world fluid dynamics experiment.

math.DS

Continuum of coupled Wasserstein gradient flows

We study a system of drift-diffusion PDEs for a potentially infinite number of incompressible phases, subject to a joint pointwise volume constraint. Our analysis is based on the interpretation as a collection of coupled Wasserstein gradient flows or, equivalently, as a gradient flow in the space of couplings under a `fibered' Wasserstein distance. We prove existence of weak solutions, long-time asymptotics, and stability with respect to the mass distribution of the phases, including the discrete to continuous limit. A key step is to establish convergence of the product of pressure gradient and density, jointly over the infinite number of phases. The underlying energy functional is the objective of entropy regularized optimal transport, which allows us to interpret the model as the relaxation of the classical Angenent-Haker-Tannenbaum (AHT) scheme to the entropic setting. However, in contrast to the AHT scheme's lack of convergence guarantees, the relaxed scheme is unconditionally convergent. We conclude with numerical illustrations of the main results.

math.AP

Domain decomposition for entropic unbalanced optimal transport

Solving large scale entropic optimal transport problems with the Sinkhorn algorithm remains challenging, and domain decomposition has been shown to be an efficient strategy for problems on large grids. Unbalanced optimal transport is a versatile variant of the balanced transport problem and its entropic regularization can be solved with an adapted Sinkhorn algorithm. However, it is a priori unclear how to apply domain decomposition to unbalanced problems since the independence of the cell problems is lost. In this article we show how this difficulty can be overcome at a theoretical and practical level and demonstrate with experiments that domain decomposition is also viable and efficient on large unbalanced entropic transport problems.

math.OC

Formulas for the $h$-mass on $1$-currents with coefficients in $\mathbb{R}^m$

We consider the minimization of the $h$-mass over normal $1$-currents in $\mathbb{R}^n$ with coefficients in $\mathbb{R}^m$ and prescribed boundary. This optimization is known as multi-material transport problem and used in the context of logistics of multiple commodities, but also as a relaxation of nonconvex optimal transport tasks such as so-called branched transport problems. The $h$-mass with norm $h$ can be defined in different ways, resulting in three functionals $\mathcal{M}_h,|\cdot|_H$, and $\mathbb{M}_h$, whose equality is the main result of this article: $\mathcal{M}_h$ is a functional on $1$-currents in the spirit of Federer and Fleming, norm $|\cdot|_H$ denotes the total variation of a Radon measure with respect to $H$ induced by $h$, and $\mathbb{M}_h$ is a mass on flat $1$-chains in the sense of Whitney. On top we introduce a new and improved notion of calibrations for the multi-material transport problem: we identify calibrations with (weak) Jacobians of optimizers of the associated convex dual problem, which yields their existence and natural regularity.

math.OC

Linearized optimal transport on manifolds

Optimal transport is a geometrically intuitive, robust and flexible metric for sample comparison in data analysis and machine learning. Its formal Riemannian structure allows for a local linearization via a tangent space approximation. This in turn leads to a reduction of computational complexity and simplifies combination with other methods that require a linear structure. Recently this approach has been extended to the unbalanced Hellinger--Kantorovich (HK) distance. In this article we further extend the framework in various ways, including measures on manifolds, the spherical HK distance, a study of the consistency of discretization via the barycentric projection, and the continuity properties of the logarithmic map for the HK distance.

math.OC

Flow updates for domain decomposition of entropic optimal transport

Domain decomposition has been shown to be a computationally efficient distributed method for solving large scale entropic optimal transport problems. However, a naive implementation of the algorithm can freeze in the limit of very fine partition cells (i.e. it asymptotically becomes stationary and does not find the global minimizer), since information can only travel slowly between cells. In practice this can be avoided by a coarse-to-fine multiscale scheme. In this article we introduce flow updates as an alternative approach. Flow updates can be interpreted as a variant of the celebrated algorithm by Angenent, Haker, and Tannenbaum, and can be combined canonically with domain decomposition. We prove convergence to the global minimizer and provide a formal discussion of its continuity limit. We give a numerical comparison with naive and multiscale domain decomposition, and show that the flow updates prevent freezing in the regime of very many cells. While the multiscale scheme is observed to be faster than the hybrid approach in general, the latter could be a viable alternative in cases where a good initial coupling is available. Our numerical experiments are based on a novel GPU implementation of domain decomposition that we describe in the appendix.

math.NA

The Riemannian geometry of Sinkhorn divergences

We propose a new metric between probability measures on a compact metric space that mirrors the Riemannian manifold-like structure of quadratic optimal transport but includes entropic regularization. Its metric tensor is given by the Hessian of the Sinkhorn divergence, a debiased variant of entropic optimal transport. We precisely identify the tangent space it induces, which turns out to be related to a Reproducing Kernel Hilbert Space (RKHS). As usual in Riemannian geometry, the distance is built by looking for shortest paths. We prove that our distance is geodesic, metrizes the weak-star topology, and is equivalent to a RKHS norm. Still it retains the geometric flavor of optimal transport: as a paradigmatic example, translations are geodesics for the quadratic cost on $\mathbb{R}^d$. We also show two negative results on the Sinkhorn divergence that may be of independent interest: that it is not jointly convex, and that its square root is not a distance because it fails to satisfy the triangle inequality.

math.OC

Transfer Operators from Batches of Unpaired Points via Entropic Transport Kernels

In this paper, we are concerned with estimating the joint probability of random variables $X$ and $Y$, given $N$ independent observation blocks $(\boldsymbol{x}^i,\boldsymbol{y}^i)$, $i=1,\ldots,N$, each of $M$ samples $(\boldsymbol{x}^i,\boldsymbol{y}^i) = \bigl((x^i_j, y^i_{σ^i(j)}) \bigr)_{j=1}^M$, where $σ^i$ denotes an unknown permutation of i.i.d. sampled pairs $(x^i_j,y_j^i)$, $j=1,\ldots,M$. This means that the internal ordering of the $M$ samples within an observation block is not known. We derive a maximum-likelihood inference functional, propose a computationally tractable approximation and analyze their properties. In particular, we prove a $Γ$-convergence result showing that we can recover the true density from empirical approximations as the number $N$ of blocks goes to infinity. Using entropic optimal transport kernels, we model a class of hypothesis spaces of density functions over which the inference functional can be minimized. This hypothesis class is particularly suited for approximate inference of transfer operators from data. We solve the resulting discrete minimization problem by a modification of the EMML algorithm to take addional transition probability constraints into account and prove the convergence of this algorithm. Proof-of-concept examples demonstrate the potential of our method.

stat.ML

A Bayesian model for dynamic mass reconstruction from PET listmode data

Positron emission tomography (PET) is a classical imaging technique to reconstruct the mass distribution of a radioactive material. If the mass distribution is static, this essentially leads to inversion of the X-ray transform. However, if the mass distribution changes temporally, the measurement signals received over time (the so-called listmode data) belong to different spatial configurations. We suggest and analyse a Bayesian approach to solve this dynamic inverse problem that is based on optimal transport regularization of the temporally changing mass distribution. Our focus lies on a rigorous derivation of the Bayesian model and the analysis of its properties, treating both the continuous as well as the discrete (finitely many detectors and time binning) setting.

math.OC

Manifold learning in Wasserstein space

This paper aims at building the theoretical foundations for manifold learning algorithms in the space of absolutely continuous probability measures $\mathcal{P}_{\mathrm{a.c.}}(\Omega)$ with $\Omega$ a compact and convex subset of $\mathbb{R}^d$, metrized with the Wasserstein-2 distance $\mathbb{W}$. We begin by introducing a construction of submanifolds $\Lambda$ in $\mathcal{P}_{\mathrm{a.c.}}(\Omega)$ equipped with metric $\mathbb{W}_\Lambda$, the geodesic restriction of $\mathbb{W}$ to $\Lambda$. In contrast to other constructions, these submanifolds are not necessarily flat, but still allow for local linearizations in a similar fashion to Riemannian submanifolds of $\mathbb{R}^d$. We then show how the latent manifold structure of $(\Lambda,\mathbb{W}_{\Lambda})$ can be learned from samples $\{\lambda_i\}_{i=1}^N$ of $\Lambda$ and pairwise extrinsic Wasserstein distances $\mathbb{W}$ on $\mathcal{P}_{\mathrm{a.c.}}(\Omega)$ only. In particular, we show that the metric space $(\Lambda,\mathbb{W}_{\Lambda})$ can be asymptotically recovered in the sense of Gromov--Wasserstein from a graph with nodes $\{\lambda_i\}_{i=1}^N$ and edge weights $W(\lambda_i,\lambda_j)$. In addition, we demonstrate how the tangent space at a sample $\lambda$ can be asymptotically recovered via spectral analysis of a suitable ``covariance operator'' using optimal transport maps from $\lambda$ to sufficiently close and diverse samples $\{\lambda_i\}_{i=1}^N$. The paper closes with some explicit constructions of submanifolds $\Lambda$ and numerical examples on the recovery of tangent spaces through spectral analysis.

stat.ML

Entropic transfer operators

We propose a new concept for the regularization and discretization of transfer and Koopman operators in dynamical systems. Our approach is based on the entropically regularized optimal transport between two probability measures. In particular, we use optimal transport plans in order to construct a finite-dimensional approximation of some transfer or Koopman operator which can be analysed computationally. We prove that the spectrum of the discretized operator converges to the one of the regularized original operator, give a detailed analysis of the relation between the discretized and the original peripheral spectrum for a rotation map on the $n$-torus and provide code for three numerical experiments, including one based on the raw trajectory data of a small biomolecule from which its dominant conformations are recovered.

math.DS

Hellinger-Kantorovich barycenter between Dirac measures

The Hellinger-Kantorovich (HK) distance is an unbalanced extension of the Wasserstein-2 distance. It was shown recently that the HK barycenter exhibits a much more complex behaviour than the Wasserstein barycenter. Motivated by this observation we study the HK barycenter in more detail for the case where the input measures are an uncountable collection of Dirac measures, in particular the dependency on the length scale parameter of HK, the question whether the HK barycenter is discrete or continuous and the relation between the expected and the empirical barycenter. The analytical results are complemented with numerical experiments that demonstrate that the HK barycenter can provide a coarse-to-fine representation of an input pointcloud or measure.

math.OC

Duality in branched transport and urban planning

In recent work arXiv:2109.07820 we have shown the equivalence of the widely used nonconvex (generalized) branched transport problem with a shape optimization problem of a street or railroad network, known as (generalized) urban planning problem. The argument was solely based on an explicit construction and characterization of competitors. In the current article we instead analyse the dual perspective associated with both problems. In more detail, the shape optimization problem involves the Wasserstein distance between two measures with respect to a metric depending on the street network. We show a Kantorovich$\unicode{x2013}$Rubinstein formula for Wasserstein distances on such street networks under mild assumptions. Further, we provide a Beckmann formulation for such Wasserstein distances under assumptions which generalize our previous result in arXiv:2109.07820. As an application we then give an alternative, duality-based proof of the equivalence of both problems under a growth condition on the transportation cost, which reveals that urban planning and branched transport can both be viewed as two bilinearly coupled convex optimization problems.

math.OC

Data-driven entropic spatially inhomogeneous evolutionary games

We introduce novel multi-agent interaction models of entropic spatially inhomogeneous evolutionary undisclosed games and their quasi-static limits. These evolutions vastly generalize first and second order dynamics. Besides the well-posedness of these novel forms of multi-agent interactions, we are concerned with the learnability of individual payoff functions from observation data. We formulate the payoff learning as a variational problem, minimizing the discrepancy between the observations and the predictions by the payoff function. The inferred payoff function can then be used to simulate further evolutions, which are fully data-driven. We prove convergence of minimizing solutions obtained from a finite number of observations to a mean field limit and the minimal value provides a quantitative error bound on the data-driven evolutions. The abstract framework is fully constructive and numerically implementable. We illustrate this on computational examples where a ground truth payoff function is known and on examples where this is not the case, including a model for pedestrian movement.

math.OC

Domain decomposition for entropy regularized optimal transport

We study Benamou's domain decomposition algorithm for optimal transport in the entropy regularized setting. The key observation is that the regularized variant converges to the globally optimal solution under very mild assumptions. We prove linear convergence of the algorithm with respect to the Kullback--Leibler divergence and illustrate the (potentially very slow) rates with numerical examples. On problems with sufficient geometric structure (such as Wasserstein distances between images) we expect much faster convergence. We then discuss important aspects of a computationally efficient implementation, such as adaptive sparsity, a coarse-to-fine scheme and parallelization, paving the way to numerically solving large-scale optimal transport problems. We demonstrate efficient numerical performance for computing the Wasserstein-2 distance between 2D images and observe that, even without parallelization, domain decomposition compares favorably to applying a single efficient implementation of the Sinkhorn algorithm in terms of runtime, memory and solution quality.

math.OC