SearcharxivSearch

arXiv subjects

Viktor Nilsson

Publications and source records attributed to Viktor Nilsson.

7 recordsLinked to original sources

Reflected Schr\"odinger Bridge Matching

Recent advances in generative modeling have enabled the efficient computation of Schr\"odinger bridges (SB) in high-dimensional settings by leveraging partially simulation-free training methods inspired by flow matching. However, these have not covered SBs with reflecting dynamics, a useful model choice with built-in guarantees that generated samples stay in the data domain. Existing alternatives for reflected SBs instead rely on more complex training based on forward--backward SDE theory, requiring expensive higher-order derivatives and sampling entire paths during training. In this article, we introduce a partially simulation-free framework that allows reflected SBs to be trained similarly to flow matching, using a new sampling method and regression target. We demonstrate our results by coupling pairs of well-known high-dimensional image datasets. Using reflected dynamics incurs negligible additional wall-clock time during both training and inference while maintaining or slightly improving generative performance.

cs.LG

A weak convergence approach to the large deviations of the dynamic Schr\"odinger problem

In this paper, we consider the large deviations for dynamical Schr\"odinger problems, using the variational approach developed by Dupuis, Ellis, Budhiraja, and others. Recent results on scaled families of Schr\"odinger problems, in particular by Bernton, Ghosal, and Nutz, and the authors, have established large deviation principles for the static problem. For the dynamic problem, only the case with a scaled Brownian motion reference process has been explored by Kato. Here, we derive large deviations results using the variational approach, with the aim of going beyond the Brownian reference dynamics considered by Kato. Specifically, we develop a uniform Laplace principle for bridge processes conditioned on their endpoints. When combined with existing results for the static problem, this leads to a large deviation principle for the corresponding (dynamic) Schr\"odinger bridge. In addition to the specific results of the paper, our work puts such large deviation questions into the weak convergence framework, and we conjecture that the results can be extended to cover also more involved types of reference dynamics. Specifically, we provide an outlook on applying the result to reflected Schr\"odinger bridges.

math.PR

Large deviations for scaled families of Schr\"odinger bridges with reflection

In this paper, we show a large deviation principle for certain sequences of static Schr\"{o}dinger bridges, typically motivated by a scale-parameter decreasing towards zero, extending existing large deviation results to cover a wider range of reference processes. Our results provide a theoretical foundation for studying convergence of such Schr\"{o}dinger bridges to their limiting optimal transport plans. Within generative modeling, Schr\"{o}dinger bridges, or entropic optimal transport problems, constitute a prominent class of methods, in part because of their computational feasibility in high-dimensional settings. Recently, Bernton et al. established a large deviation principle, in the small-noise limit, for fixed-cost entropic optimal transport problems. In this paper, we address an open problem posed by Bernton et al. and extend their results to hold for Schr\"{o}dinger bridges associated with certain sequences of more general reference measures with enough regularity in a similar small-noise limit. These can be viewed as sequences of entropic optimal transport plans with non-fixed cost functions. Using a detailed analysis of the associated Skorokhod maps and transition densities, we show that the new large deviation results cover Schr\"{o}dinger bridges where the reference process is a reflected diffusion on bounded convex domains, corresponding to recently introduced model choices in the generative modeling literature.

math.PR

Efficient Flow Matching using Latent Variables

Flow matching models have shown great potential in image generation tasks among probabilistic generative models. However, most flow matching models in the literature do not explicitly utilize the underlying clustering structure in the target data when learning the flow from a simple source distribution like the standard Gaussian. This leads to inefficient learning, especially for many high-dimensional real-world datasets, which often reside in a low-dimensional manifold. To this end, we present $\texttt{Latent-CFM}$, which provides efficient training strategies by conditioning on the features extracted from data using pretrained deep latent variable models. Through experiments on synthetic data from multi-modal distributions and widely used image benchmark datasets, we show that $\texttt{Latent-CFM}$ exhibits improved generation quality with significantly less training and computation than state-of-the-art flow matching models by adopting pretrained lightweight latent variable models. Beyond natural images, we consider generative modeling of spatial fields stemming from physical processes. Using a 2d Darcy flow dataset, we demonstrate that our approach generates more physically accurate samples than competing approaches. In addition, through latent space analysis, we demonstrate that our approach can be used for conditional image generation conditioned on latent features, which adds interpretability to the generation process.

cs.CV

REMEDI: Corrective Transformations for Improved Neural Entropy Estimation

Information theoretic quantities play a central role in machine learning. The recent surge in the complexity of data and models has increased the demand for accurate estimation of these quantities. However, as the dimension grows the estimation presents significant challenges, with existing methods struggling already in relatively low dimensions. To address this issue, in this work, we introduce $\texttt{REMEDI}$ for efficient and accurate estimation of differential entropy, a fundamental information theoretic quantity. The approach combines the minimization of the cross-entropy for simple, adaptive base models and the estimation of their deviation, in terms of the relative entropy, from the data density. Our approach demonstrates improvement across a broad spectrum of estimation tasks, encompassing entropy estimation on both synthetic and natural data. Further, we extend important theoretical consistency results to a more generalized setting required by our approach. We illustrate how the framework can be naturally extended to information theoretic supervised learning models, with a specific focus on the Information Bottleneck approach. It is demonstrated that the method delivers better accuracy compared to the existing methods in Information Bottleneck. In addition, we explore a natural connection between $\texttt{REMEDI}$ and generative modeling using rejection sampling and Langevin dynamics.

stat.ML

Large deviations for interacting particle dynamics for finding mixed equilibria in zero-sum games

Finding equilibrium points in continuous minmax games has become a key problem within machine learning, in part due to its connection to the training of generative adversarial networks and reinforcement learning. Because of existence and robustness issues, recent developments have shifted from pure equilibria to focusing on mixed equilibrium points. In this work we consider a method for finding mixed equilibria in two-layer zero-sum games based on entropic regularisation, where the two competing strategies are represented by two sets of interacting particles. We show that the sequence of empirical measures of the particle system satisfies a large deviation principle as the number of particles grows to infinity, and how this implies convergence of the empirical measure and the associated Nikaid\^o-Isoda error, complementing existing law of large numbers results.

stat.ML

Probabilistic dose prediction using mixture density networks for automated radiation therapy treatment planning

We demonstrate the application of mixture density networks (MDNs) in the context of automated radiation therapy treatment planning. It is shown that an MDN can produce good predictions of dose distributions as well as reflect uncertain decision making associated with inherently conflicting clinical tradeoffs, in contrast to deterministic methods previously investigated in literature. A two-component Gaussian MDN is trained on a set of treatment plans for postoperative prostate patients with varying extents to which rectum dose sparing was prioritized over target coverage. Examination on a test set of patients shows that the predicted modes follow their respective ground truths well both spatially and in terms of their dose-volume histograms. A special dose mimicking method based on the MDN output is used to produce deliverable plans and thereby showcase the usability of voxel-wise predictive densities. Thus, this type of MDN may serve to support clinicians in managing clinical tradeoffs and has the potential to improve quality of plans produced by an automated treatment planning pipeline.

physics.med-ph