SearcharxivSearch

arXiv subjects

Pierre Nyquist

Publications and source records attributed to Pierre Nyquist.

At least 19 recordsLinked to original sources

Reflected Schr\"odinger Bridge Matching

Recent advances in generative modeling have enabled the efficient computation of Schr\"odinger bridges (SB) in high-dimensional settings by leveraging partially simulation-free training methods inspired by flow matching. However, these have not covered SBs with reflecting dynamics, a useful model choice with built-in guarantees that generated samples stay in the data domain. Existing alternatives for reflected SBs instead rely on more complex training based on forward--backward SDE theory, requiring expensive higher-order derivatives and sampling entire paths during training. In this article, we introduce a partially simulation-free framework that allows reflected SBs to be trained similarly to flow matching, using a new sampling method and regression target. We demonstrate our results by coupling pairs of well-known high-dimensional image datasets. Using reflected dynamics incurs negligible additional wall-clock time during both training and inference while maintaining or slightly improving generative performance.

cs.LG

A weak convergence approach to the large deviations of the dynamic Schr\"odinger problem

In this paper, we consider the large deviations for dynamical Schr\"odinger problems, using the variational approach developed by Dupuis, Ellis, Budhiraja, and others. Recent results on scaled families of Schr\"odinger problems, in particular by Bernton, Ghosal, and Nutz, and the authors, have established large deviation principles for the static problem. For the dynamic problem, only the case with a scaled Brownian motion reference process has been explored by Kato. Here, we derive large deviations results using the variational approach, with the aim of going beyond the Brownian reference dynamics considered by Kato. Specifically, we develop a uniform Laplace principle for bridge processes conditioned on their endpoints. When combined with existing results for the static problem, this leads to a large deviation principle for the corresponding (dynamic) Schr\"odinger bridge. In addition to the specific results of the paper, our work puts such large deviation questions into the weak convergence framework, and we conjecture that the results can be extended to cover also more involved types of reference dynamics. Specifically, we provide an outlook on applying the result to reflected Schr\"odinger bridges.

math.PR

Large deviations for scaled families of Schr\"odinger bridges with reflection

In this paper, we show a large deviation principle for certain sequences of static Schr\"{o}dinger bridges, typically motivated by a scale-parameter decreasing towards zero, extending existing large deviation results to cover a wider range of reference processes. Our results provide a theoretical foundation for studying convergence of such Schr\"{o}dinger bridges to their limiting optimal transport plans. Within generative modeling, Schr\"{o}dinger bridges, or entropic optimal transport problems, constitute a prominent class of methods, in part because of their computational feasibility in high-dimensional settings. Recently, Bernton et al. established a large deviation principle, in the small-noise limit, for fixed-cost entropic optimal transport problems. In this paper, we address an open problem posed by Bernton et al. and extend their results to hold for Schr\"{o}dinger bridges associated with certain sequences of more general reference measures with enough regularity in a similar small-noise limit. These can be viewed as sequences of entropic optimal transport plans with non-fixed cost functions. Using a detailed analysis of the associated Skorokhod maps and transition densities, we show that the new large deviation results cover Schr\"{o}dinger bridges where the reference process is a reflected diffusion on bounded convex domains, corresponding to recently introduced model choices in the generative modeling literature.

math.PR

A weak convergence approach to large deviations for stochastic approximations

The theory of stochastic approximations form the theoretical foundation for studying convergence properties of many popular recursive learning algorithms in statistics, machine learning and statistical physics. Large deviations for stochastic approximations provide asymptotic estimates of the probability that the learning algorithm deviates from its expected path, given by a limit ODE, and the large deviation rate function gives insights to the most likely way that such deviations occur. In this paper we prove a large deviation principle for general stochastic approximations with state-dependent Markovian noise and decreasing step size. Using the weak convergence approach to large deviations, we generalize previous results for stochastic approximations and identify the appropriate scaling sequence for the large deviation principle. We also give a new representation for the rate function, in which the rate function is expressed as an action functional involving the family of Markov transition kernels. Examples of learning algorithms that are covered by the large deviation principle include stochastic gradient descent, persistent contrastive divergence and the Wang-Landau algorithm.

math.PR

Large deviations for Independent Metropolis Hastings and Metropolis-adjusted Langevin algorithm

In this paper, we prove large deviation principles for the empirical measures associated with the Independent Metropolis Hastings (IMH) sampler and the Metropolis-adjusted Langevin Algorithm (MALA). These are the first large deviation results for empirical measures of Markov chains arising from specific Metropolis-Hastings methods on a continuous state space. Moreover, we show that the existing large deviation framework, that we developed in a previous work (Milinanni and Nyquist, 2024), does not cover the Random Walk Metropolis sampler, even in cases when the underlying Markov chain is geometrically ergodic.

math.PR

REMEDI: Corrective Transformations for Improved Neural Entropy Estimation

Information theoretic quantities play a central role in machine learning. The recent surge in the complexity of data and models has increased the demand for accurate estimation of these quantities. However, as the dimension grows the estimation presents significant challenges, with existing methods struggling already in relatively low dimensions. To address this issue, in this work, we introduce $\texttt{REMEDI}$ for efficient and accurate estimation of differential entropy, a fundamental information theoretic quantity. The approach combines the minimization of the cross-entropy for simple, adaptive base models and the estimation of their deviation, in terms of the relative entropy, from the data density. Our approach demonstrates improvement across a broad spectrum of estimation tasks, encompassing entropy estimation on both synthetic and natural data. Further, we extend important theoretical consistency results to a more generalized setting required by our approach. We illustrate how the framework can be naturally extended to information theoretic supervised learning models, with a specific focus on the Information Bottleneck approach. It is demonstrated that the method delivers better accuracy compared to the existing methods in Information Bottleneck. In addition, we explore a natural connection between $\texttt{REMEDI}$ and generative modeling using rejection sampling and Langevin dynamics.

stat.ML

UQSA -- An R-Package for Uncertainty Quantification and Sensitivity Analysis for Biochemical Reaction Network Models

Biochemical reaction models describing subcellular processes generally come with a large uncertainty. To be able to account for this during the modeling process, we have developed the R-package UQSA, performing uncertainty quantification and sensitivity analysis in an integrated fashion. UQSA is designed for fast sampling of complicated multi-dimensional parameter distributions, using efficient Markov chain Monte Carlo (MCMC) sampling techniques and Vine-copulas to model complicated joint distributions. We perform MCMC sampling both from stochastic and deterministic models, in either likelihood-free or likelihood-based settings. In the likelihood-free case, we use Approximate Bayesian Computation (ABC), while for likelihood-based sampling we provide different algorithms, including the fast geometry-informed algorithm SMMALA (Simplified Manifold Metropolis-Adjusted Langevin Algorithm). The uncertainty quantification can be followed by a variance decomposition-based global sensitivity analysis. We are aiming for biochemical models, but UQSA can be used for any type of reaction networks. The use of Vine-copulas allows us to describe, evaluate, and sample from complicated parameter distributions, as well as adding new datasets in a sequential manner without redoing the previous parameter fit. The code is written in R, with C as a backend to improve speed. We use the SBtab table format for Systems Biology projects for the model description as well as the experimental data. An event system allows the user to model complicated transient input, common within, e.g., neuroscience. UQSA has an extensive documentation with several examples describing different types of models and data. The code has been tested on up to 2000 cores on several nodes on a computing cluster, but we also include smaller examples that can be run on a laptop. Source code: https://github.com/icpm-kth/uqsa

q-bio.QM

A large deviation principle for the empirical measures of Metropolis-Hastings chains

To sample from a given target distribution, Markov chain Monte Carlo (MCMC) sampling relies on constructing an ergodic Markov chain with the target distribution as its invariant measure. For any MCMC method, an important question is how to evaluate its efficiency. One approach is to consider the associated empirical measure and how fast it converges to the stationary distribution of the underlying Markov process. Recently, this question has been considered from the perspective of large deviation theory, for different types of MCMC methods, including, e.g., non-reversible Metropolis-Hastings on a finite state space, non-reversible Langevin samplers, the zig-zag sampler, and parallell tempering. This approach, based on large deviations, has proven successful in analysing existing methods and designing new, efficient ones. However, for the Metropolis-Hastings algorithm on more general state spaces, the workhorse of MCMC sampling, the same techniques have not been available for analysing performance, as the underlying Markov chain dynamics violate the conditions used to prove existing large deviation results for empirical measures of a Markov chain. This also extends to methods built on the same idea as Metropolis-Hastings, such as the Metropolis-Adjusted Langevin Method or ABC-MCMC. In this paper, we take the first steps towards such a large-deviations based analysis of Metropolis-Hastings-like methods, by proving a large deviation principle for the the empirical measures of Metropolis-Hastings chains. In addition, we characterize the rate function and its properties in terms of the acceptance- and rejection-part of the Metropolis-Hastings dynamics.

math.PR

Large deviations for interacting particle dynamics for finding mixed equilibria in zero-sum games

Finding equilibrium points in continuous minmax games has become a key problem within machine learning, in part due to its connection to the training of generative adversarial networks and reinforcement learning. Because of existence and robustness issues, recent developments have shifted from pure equilibria to focusing on mixed equilibrium points. In this work we consider a method for finding mixed equilibria in two-layer zero-sum games based on entropic regularisation, where the two competing strategies are represented by two sets of interacting particles. We show that the sequence of empirical measures of the particle system satisfies a large deviation principle as the number of particles grows to infinity, and how this implies convergence of the empirical measure and the associated Nikaid\^o-Isoda error, complementing existing law of large numbers results.

stat.ML

Sensitivity Approximation by the Peano-Baker Series

In this paper we develop a new method for numerically approximating sensitivities in parameter-dependent ordinary differential equations (ODEs). Our approach, intended for situations where the standard forward and adjoint sensitivity analyses become too computationally costly for practical purposes, is based on the Peano-Baker series from control theory. Using this series, we construct a representation of the sensitivity matrix $\mathbf{S}$ and, from this representation, a numerical method for approximating $\mathbf{S}$. We prove that, under standard regularity assumptions, the error of our method scales as $\mathcal{O}(\Delta t^2_{\text{max}})$, where $\Delta t_{\text{max}}$ is the largest time step used when numerically solving the ODE. We illustrate the performance of the method in several numerical experiments, taken from both the systems biology setting and more classical dynamical systems. The experiments show the sought-after improvement in running time of our method compared to the forward sensitivity approach. In experiments involving a random linear system, the forward approach requires roughly $\sqrt{n}$ longer computational time, where $n$ is the dimension of the parameter space, than our proposed method.

math.NA

Quasistationary Distributions and Ergodic Control Problems

We introduce and study the basic properties of two ergodic stochastic control problems associated with the quasistationary distribution (QSD) of a diffusion process $X$ relative to a bounded domain. The two problems are in some sense dual, with one defined in terms of the generator associated with $X$ and the other in terms of its adjoint. Besides proving wellposedness of the associated Hamilton-Jacobi-Bellman equations, we describe how they can be used to characterize important properties of the QSD. Of particular note is that the QSD itself can be identified, up to normalization, in terms of the cost potential of the control problem associated with the adjoint.

math.OC

Large deviations and gradient flows for the Brownian one-dimensional hard-rod system

We study a system of hard rods of finite size in one space dimension, which move by Brownian noise while avoiding overlap. We consider a scaling in which the number of particles tends to infinity while the volume fraction of the rods remains constant; in this limit the empirical measure of the rod positions converges almost surely to a deterministic limit evolution. We prove a large-deviation principle on path space for the empirical measure, by exploiting a one-to-one mapping between the hard-rod system and a system of non-interacting particles on a shorter domain. The large-deviation principle naturally identifies a gradient-flow structure for the limit evolution, with clear interpretations for both the driving functional (an `entropy') and the dissipation, which in this case is the Wasserstein dissipation. This study is inspired by recent developments in the continuum modelling of multiple-species interacting particle systems with finite-size effects; for such systems many different modelling choices appear in the literature, raising the question how one can understand such choices in terms of more microscopic models. The results of this paper give a clear answer to this question, albeit for the simpler one-dimensional hard-rod system. For this specific system this result provides a clear understanding of the value and interpretation of different modelling choices, while giving hints for more general systems.

math-ph

Large deviations for the empirical measure of the zig-zag process

The zig-zag process is a piecewise deterministic Markov process in position and velocity space. The process can be designed to have an arbitrary Gibbs type marginal probability density for its position coordinate, which makes it suitable for Monte Carlo simulation of continuous probability distributions. An important question in assessing the efficiency of this method is how fast the empirical measure converges to the stationary distribution of the process. In this paper we provide a partial answer to this question by characterizing the large deviations of the empirical measure from the stationary distribution. Based on the Feng-Kurtz approach, we develop an abstract framework aimed at encompassing piecewise deterministic Markov processes in position-velocity space. We derive explicit conditions for the zig-zag process to allow the Donsker-Varadhan variational formulation of the rate function, both for a compact setting (the torus) and one-dimensional Euclidean space. Finally we derive an explicit expression for the Donsker-Varadhan functional for the case of a compact state space and use this form of the rate function to address a key question concerning the optimal choice of the switching rate of the zig-zag process.

math.PR

Importance sampling for a simple Markovian intensity model using subsolutions

This paper considers importance sampling for estimation of rare-event probabilities in a specific collection of Markovian jump processes used for e.g. modelling of credit risk. Previous attempts at designing importance sampling algorithms have resulted in poor performance and the main contribution of the paper is the design of efficient importance sampling algorithms using subsolutions. The dynamics of the jump processes causes the corresponding Hamilton-Jacobi equations to have an intricate state-dependence, which makes the design of efficient algorithms difficult. We provide theoretical results that quantify the performance of importance sampling algorithms in general and construct asymptotically optimal algorithms for some examples. The computational gain compared to standard Monte Carlo is illustrated by numerical examples.

math.PR

A large deviations analysis of certain qualitative properties of parallel tempering and infinite swapping algorithms

Parallel tempering, or replica exchange, is a popular method for simulating complex systems. The idea is to run parallel simulations at different temperatures, and at a given swap rate exchange configurations between the parallel simulations. From the perspective of large deviations it is optimal to let the swap rate tend to infinity and it is possible to construct a corresponding simulation scheme, known as infinite swapping. In this paper we propose a novel use of large deviations for empirical measures for a more detailed analysis of the infinite swapping limit in the setting of continuous time jump Markov processes. Using the large deviations rate function and associated stochastic control problems we consider a diagnostic based on temperature assignments, which can be easily computed during a simulation. We show that the convergence of this diagnostic to its a priori known limit is a necessary condition for the convergence of infinite swapping. The rate function is also used to investigate the impact of asymmetries in the underlying potential landscape, and where in the state space poor sampling is most likely to occur.

math.PR

Large deviations for weighted empirical measures arising in importance sampling

Importance sampling is a popular method for efficient computation of various properties of a distribution such as probabilities, expectations, quantiles etc. The output of an importance sampling algorithm can be represented as a weighted empirical measure, where the weights are given by the likelihood ratio between the original distribution and the sampling distribution. In this paper the efficiency of an importance sampling algorithm is studied by means of large deviations for the weighted empirical measure. The main result, which is stated as a Laplace principle for the weighted empirical measure arising in importance sampling, can be viewed as a weighted version of Sanov's theorem. The main theorem is applied to quantify the performance of an importance sampling algorithm over a collection of subsets of a given target set as well as quantile estimates. The analysis yields an estimate of the sample size needed to reach a desired precision as well as of the reduction in cost for importance sampling compared to standard Monte Carlo.

math.PR

Large deviations for multidimensional state-dependent shot noise processes

Shot noise processes are used in applied probability to model a variety of physical systems in, for example, teletraffic theory, insurance and risk theory and in the engineering sciences. In this work we prove a large deviation principle for the sample-paths of a general class of multidimensional state-dependent Poisson shot noise processes. The result covers previously known large deviation results for one dimensional state-independent shot noise processes with light tails. We use the weak convergence approach to large deviations, which reduces the proof to establishing the appropriate convergence of certain controlled versions of the original processes together with relevant results on existence and uniqueness.

math.PR

Min-max representations of viscosity solutions of Hamilton-Jacobi equations and applications in rare-event simulation

In this paper a duality relation between the Mañé potential and Mather's action functional is derived in the context of convex and state-dependent Hamiltonians. The duality relation is used to obtain min-max representations of viscosity solutions of first order Hamilton-Jacobi equations. These min-max representations naturally suggest class\-es of subsolutions of Hamilton-Jacobi equations that arise in the theory of large deviations. The subsolutions, in turn, are good candidates for designing efficient rare-event simulation algorithms.

math.AP