SearcharxivSearch

arXiv subjects

Lukas Trottner

Publications and source records attributed to Lukas Trottner.

16 recordsLinked to original sources

Statistical Convergence of Spherical First Hitting Diffusion Models

Denoising diffusion models have evolved into a state-of-the-art method for tasks in various fields, such as denoising and generation of images, text generation, or generation of synthetic data for training of other machine learning models. First hitting diffusion models (FHDM) are a particular class of denoising diffusion models with \textit{random} adaptive generation time tailored to generate data on a known manifold. Building on the conditioning framework of Doob's $h$-transform these models leverage the given information on the target data manifold to demonstrate strong performance across tasks while offering distinct features such as time-homogeneous dynamics of the generating process and a reduced average simulation time. Even though the theoretical investigation of standard forward-backward diffusion models has attracted much attention in the recent past, the statistical convergence properties of FHDMs are not yet understood. In this work, we show that, up to logarithmic factors, FHDMs achieve the minimax optimal convergence rate in total variation for spherically supported Sobolev smooth data distributions. In particular, this is the first statistical optimality result for denoising diffusion modelling with random generation time.

math.ST

Reflected diffusion models adapt to low-dimensional data

While the mathematical foundations of score-based generative models are increasingly well understood for unconstrained Euclidean spaces, many practical applications involve data restricted to bounded domains. This paper provides a statistical analysis of reflected diffusion models on the hypercube $[0,1]^D$ for target distributions supported on $d$-dimensional linear subspaces. A primary challenge in this setting is the absence of Gaussian transition kernels, which play a central role in standard theory in $\mathbb{R}^D$. By employing an easily implementable infinite series expansion of the transition densities, we develop analytic tools to bound the score function and its approximation by sparse ReLU networks. For target densities with Sobolev smoothness $α$, we establish a convergence rate in the $1$-Wasserstein distance of order $n^{-\frac{α+1-δ}{2α+d}}$ for arbitrarily small $δ> 0$, demonstrating that the generative algorithm fully adapts to the intrinsic dimension $d$. These results confirm that the presence of reflecting boundaries does not degrade the fundamental statistical efficiency of the diffusion paradigm, matching the almost optimal rates known for unconstrained settings.

math.ST

In situ Learning-Based Spin Engineering of Pulsed Dynamic Nuclear Polarization

Pulsed Dynamic Nuclear Polarization (DNP) is currently receiving substantial interest as a means to enhance the sensitivity of nuclear magnetic resonance (NMR) and magnetic resonance imaging (MRI) by orders of magnitude. It has also received much attention as a central ingredient in many modalities of electron spin-involved quantum sensing. Relative to spin engineering associated with NMR, the design of efficient pulsed DNP experiments with a broad experimental scope are challenged by large electron-nuclear spin systems, large electron spin-involved interactions, and instrumental non-idealities and limitations. All of this may challenge traditional NMR-like theoretical and numerical pulse sequence engineering. Exploiting state-of-the-art instrumentation and taking advantage of the high sensitivity of DNP relative to NMR, we here demonstrate the use of combinations of Bayesian machine learning methods and constrained random walk procedures to design pulse sequences \textit{in situ}, by experiments, directly on the spin systems responding to spectrometer instructions. For trityl and nitroxide samples, it is demonstrated that efficient broadband DNP pulse sequences can be designed in situ with experimental protocols benchmarked against in silico analogs.

physics.chem-ph

Statistical guarantees for denoising reflected diffusion models

In recent years, denoising diffusion models have become a crucial area of research due to their abundance in the rapidly expanding field of generative AI. While recent statistical advances have delivered explanations for the generation ability of idealised denoising diffusion models for high-dimensional target data, implementations introduce thresholding procedures for the generating process to overcome issues arising from the unbounded state space of such models. This mismatch between theoretical design and implementation of diffusion models has been addressed empirically by using a \emph{reflected} diffusion process as the driver of noise instead. In this paper, we study statistical guarantees of these denoising reflected diffusion models. In particular, under Sobolev smoothness assumptions, we establish rates of convergence in total variation which, up to a polylogarithmic factor, match the minimax lower bound. Our main contributions include the statistical analysis of this novel class of denoising reflected diffusion models and a refined score approximation method in both time and space, leveraging spectral decomposition and rigorous neural network analysis.

math.ST

Beyond Fixed Horizons: A Theoretical Framework for Adaptive Denoising Diffusions

We introduce a new class of generative diffusion models that, unlike conventional denoising diffusion models, achieve a time-homogeneous structure for both the noising and denoising processes, allowing the number of steps to adaptively adjust based on the noise level. This is accomplished by conditioning the forward process using Doob's $h$-transform, which terminates the process at a suitable sampling distribution at a random time. The model is particularly well suited for generating data with lower intrinsic dimensions, as the termination criterion simplifies to a first-hitting rule. A key feature of the model is its adaptability to the target data, enabling a variety of downstream tasks using a pre-trained unconditional generative model. These tasks include natural conditioning through appropriate initialisation of the denoising process and classification of noisy data.

stat.ML

Model-free filtering in high dimensions via projection and score-based diffusions

We consider the problem of recovering a latent signal $X$ from its noisy observation $Y$. The unknown law $\mathbb{P}^X$ of $X$, and in particular its support $\mathscr{M}$, are accessible only through a large sample of i.i.d.\ observations. We further assume $\mathscr{M}$ to be a low-dimensional submanifold of a high-dimensional Euclidean space $\mathbb{R}^d$. As a filter or denoiser $\widehat X$, we suggest an estimator of the metric projection $π_{\mathscr{M}}(Y)$ of $Y$ onto the manifold $\mathscr{M}$. To compute this estimator, we study an auxiliary semiparametric model in which $Y$ is obtained by adding isotropic Laplace noise to $X$. Using score matching within a corresponding diffusion model, we obtain an estimator of the Bayesian posterior $\mathbb{P}^{X \mid Y}$ in this setup. Our main theoretical results show that, in the limit of high dimension $d$, this posterior $\mathbb{P}^{X\mid Y}$ is concentrated near the desired metric projection $π_{\mathscr{M}}(Y)$.

math.ST

Multivariate change estimation for a stochastic heat equation from local measurements

We study a stochastic heat equation with piecewise constant diffusivity $θ$ having a jump at a hypersurface $Γ$ that splits the underlying space $[0,1]^d$, $d\geq2,$ into two disjoint sets $Λ_-\cupΛ_+.$ Based on multiple spatially localized measurement observations on a regular $δ$-grid of $[0,1]^d$, we propose a joint M-estimator for the diffusivity values and the set $Λ_+$ that is inspired by statistical image reconstruction methods. We study convergence of the domain estimator $\hatΛ_+$ in the vanishing resolution level regime $δ\to 0$ and with respect to the expected symmetric difference pseudometric. As a first main finding we give a characterization of the convergence rate for $\hatΛ_+$ in terms of the complexity of $Γ$ measured by the number of intersecting hypercubes from the regular $δ$-grid. Furthermore, for the special case of domains $Λ_+$ that are built from hypercubes from the $δ$-grid, we demonstrate that perfect identification with overwhelming probability is possible with a slight modification of the estimation approach. Implications of our general results are discussed under two specific structural assumptions on $Λ_+$. For a $β$-Hölder smooth boundary fragment $Γ$, the set $Λ_+$ is estimated with rate $δ^β$. If we assume $Λ_+$ to be convex, we obtain a $δ$-rate. While our approach only aims at optimal domain estimation rates, we also demonstrate consistency of our diffusivity estimators, which is strengthened to a CLT at minimax optimal rate for sets $Λ_+$ anchored on the $δ$-grid.

math.ST

Change point estimation for a stochastic heat equation

We study a change point model based on a stochastic partial differential equation (SPDE) corresponding to the heat equation governed by the weighted Laplacian $Δ_\vartheta = \nabla\vartheta\nabla$, where $\vartheta=\vartheta(x)$ is a space-dependent diffusivity. As a basic problem the domain $(0,1)$ is considered with a piecewise constant diffusivity with a jump at an unknown point $τ$. Based on local measurements of the solution in space with resolution $δ$ over a finite time horizon, we construct a simultaneous M-estimator for the diffusivity values and the change point. The change point estimator converges at rate $δ$, while the diffusivity constants can be recovered with convergence rate $δ^{3/2}$. Moreover, when the diffusivity parameters are known and the jump height vanishes with the spatial resolution tending to zero, we derive a limit theorem for the change point estimator and identify the limiting distribution. For the mathematical analysis, a precise understanding of the SPDE with discontinuous $\vartheta$, tight concentration bounds for quadratic functionals in the solution, and a generalisation of classical M-estimators are developed.

math.ST

Data-driven rules for multidimensional reflection problems

Over the recent past data-driven algorithms for solving stochastic optimal control problems in face of model uncertainty have become an increasingly active area of research. However, for singular controls and underlying diffusion dynamics the analysis has so far been restricted to the scalar case. In this paper we fill this gap by studying a multivariate singular control problem for reversible diffusions with controls of reflection type. Our contributions are threefold. We first explicitly determine the long-run average costs as a domain-dependent functional, showing that the control problem can be equivalently characterized as a shape optimization problem. For given diffusion dynamics, assuming the optimal domain to be strongly star-shaped, we then propose a gradient descent algorithm based on polytope approximations to numerically determine a cost-minimizing domain. Finally, we investigate data-driven solutions when the diffusion dynamics are unknown to the controller. Using techniques from nonparametric statistics for stochastic processes, we construct an optimal domain estimator, whose static regret is bounded by the minimax optimal estimation rate of the unreflected process' invariant density. In the most challenging situation, when the dynamics must be learned simultaneously to controlling the process, we develop an episodic learning algorithm to overcome the emerging exploration-exploitation dilemma and show that given the static regret as a baseline, the loss in its sublinear regret per time unit is of natural order compared to the one-dimensional case.

math.OC

Markov additive friendships

The Wiener--Hopf factorisation of a Lévy or Markov additive process describes the way that it attains new maxima and minima in terms of a pair of so-called ladder height processes. Vigon's theory of friendship for Lévy processes addresses the inverse problem: when does a process exist which has certain prescribed ladder height processes? We give a complete answer to this problem for Markov additive processes, provide simpler sufficient conditions for constructing processes using friendship, and address in part the question of the uniqueness of the Wiener--Hopf factorisation for Markov additive processes.

math.PR

Covariate shift in nonparametric regression with Markovian design

Covariate shift in regression problems and the associated distribution mismatch between training and test data is a commonly encountered phenomenon in machine learning. In this paper, we extend recent results on nonparametric convergence rates for i.i.d. data to Markovian dependence structures. We demonstrate that under Hölder smoothness assumptions on the regression function, convergence rates for the generalization risk of a Nadaraya-Watson kernel estimator are determined by the similarity between the invariant distributions associated to source and target Markov chains. The similarity is explicitly captured in terms of a bandwidth-dependent similarity measure recently introduced in Pathak, Ma and Wainwright [ICML, 2022]. Precise convergence rates are derived for the particular cases of finite Markov chains and spectral gap Markov chains for which the similarity measure between their invariant distributions grows polynomially with decreasing bandwidth. For the latter, we extend the notion of a distribution transfer exponent from Kpotufe and Martinet [Ann. Stat., 49(6), 2021] to kernel transfer exponents of uniformly ergodic Markov chains in order to generate a rich class of Markov kernel pairs for which convergence guarantees for the covariate shift problem can be formulated.

math.ST

Concentration analysis of multivariate elliptic diffusion processes

We prove concentration inequalities and associated PAC bounds for continuous- and discrete-time additive functionals for possibly unbounded functions of multivariate, nonreversible diffusion processes. Our analysis relies on an approach via the Poisson equation allowing us to consider a very broad class of subexponentially ergodic processes. These results add to existing concentration inequalities for additive functionals of diffusion processes which have so far been only available for either bounded functions or for unbounded functions of processes from a significantly smaller class. We demonstrate the power of these exponential inequalities by two examples of very different areas. Considering a possibly high-dimensional parametric nonlinear drift model under sparsity constraints, we apply the continuous-time concentration results to validate the restricted eigenvalue condition for Lasso estimation, which is fundamental for the derivation of oracle inequalities. The results for discrete additive functionals are used to investigate the unadjusted Langevin MCMC algorithm for sampling of moderately heavy-tailed densities $π$. In particular, we provide PAC bounds for the sample Monte Carlo estimator of integrals $π(f)$ for polynomially growing functions $f$ that quantify sufficient sample and step sizes for approximation within a prescribed margin with high probability.

math.PR

Stability of overshoots of Markov additive processes

We prove precise stability results for overshoots of Markov additive processes (MAPs) with finite modulating space. Our approach is based on the Markovian nature of overshoots of MAPs whose mixing and ergodic properties are investigated in terms of the characteristics of the MAP. On our way we extend fluctuation theory of MAPs, contributing among others to the understanding of the Wiener-Hopf factorization for MAPs by generalizing Vigon's équations amicales inversés known for Lévy processes. Using the Lamperti transformation the results can be applied to self-similar Markov processes. Among many possible applications, we study the mixing behavior of stable processes sampled at first hitting times as a concrete example.

math.PR

Mixing it up: A general framework for Markovian statistics

Up to now, the nonparametric analysis of multidimensional continuous-time Markov processes has focussed strongly on specific model choices, mostly related to symmetry of the semigroup. While this approach allows to study the performance of estimators for the characteristics of the process in the minimax sense, it restricts the applicability of results to a rather constrained set of stochastic processes and in particular hardly allows incorporating jump structures. As a consequence, for many models of applied and theoretical interest, no statement can be made about the robustness of typical statistical procedures beyond the beautiful, but limited framework available in the literature. To close this gap, we identify $β$-mixing of the process and heat kernel bounds on the transition density as a suitable combination to obtain $\sup$-norm and $L^2$ kernel invariant density estimation rates matching the case of reversible multidimenisonal diffusion processes and outperforming density estimation based on discrete i.i.d. or weakly dependent data. Moreover, we demonstrate how up to $\log$-terms, optimal $\sup$-norm adaptive invariant density estimation can be achieved within our general framework based on tight uniform moment bounds and deviation inequalities for empirical processes associated to additive functionals of Markov processes. The underlying assumptions are verifiable with classical tools from stability theory of continuous time Markov processes and PDE techniques, which opens the door to evaluate statistical performance for a vast amount of Markov models. We highlight this point by showing how multidimensional jump SDEs with Lévy driven jump part under different coefficient assumptions can be seamlessly integrated into our framework, thus establishing novel adaptive $\sup$-norm estimation rates for this class of processes.

math.ST

Learning to reflect: A unifying approach for data-driven stochastic control strategies

Stochastic optimal control problems have a long tradition in applied probability, with the questions addressed being of high relevance in a multitude of fields. Even though theoretical solutions are well understood in many scenarios, their practicability suffers from the assumption of known dynamics of the underlying stochastic process, raising the statistical challenge of developing purely data-driven strategies. For the mathematically separated classes of continuous diffusion processes and Lévy processes, we show that developing efficient strategies for related singular stochastic control problems can essentially be reduced to finding rate-optimal estimators with respect to the sup-norm risk of objects associated to the invariant distribution of ergodic processes which determine the theoretical solution of the control problem. From a statistical perspective, we exploit the exponential $β$-mixing property as the common factor of both scenarios to drive the convergence analysis, indicating that relying on general stability properties of Markov processes is a sufficiently powerful and flexible approach to treat complex applications requiring statistical methods. We show moreover that in the Lévy case $-$ even though per se jump processes are more difficult to handle both in statistics and control theory $-$ a fully data-driven strategy with regret of significantly better order than in the diffusion case can be constructed.

math.ST