SearcharxivSearch

arXiv subjects

Lingjiong Zhu

Publications and source records attributed to Lingjiong Zhu.

At least 19 recordsLinked to original sources

Sharp hypocoercive convergence estimates for underdamped Langevin dynamics with specular reflection

We study the underdamped (kinetic) Langevin dynamics confined to a bounded domain $Ω\subset\mathbb{R}^d$ by specular reflection of the velocity at the boundary. This process is the natural momentum-based analogue of the normally reflected overdamped Langevin diffusion, and it is used in practice for constrained sampling. Assuming only that the position marginal $μ_x\propto e^{-U}$ satisfies a Poincaré inequality on $Ω$ with constant $m>0$, $\nabla^2U\succeq-K\,\mathrm{Id}$ and $Ω$ is a convex domain, we prove that the law converges to the Gibbs measure exponentially fast in $L^2$, with an explicit rate that scales like $\sqrt m$, which is optimal when $U$ is convex. Since the normally reflected overdamped dynamics converges exactly at rate $m$, this establishes a square-root acceleration for constrained sampling in the small-gap regime when $m$ is small, matching the acceleration known in the unconstrained case. The explicit rate is the same in the unconstrained setting of Fan--Li--Lu. The proof adapts the modified $L^2$ hypocoercivity method of Dolbeault--Mouhot--Schmeiser with the gap-shifted corrector of Fan--Li--Lu. The specular symmetry makes the transport operator antisymmetric, and that the corrector automatically selects the Neumann realization of the overdamped generator, which is precisely the boundary condition that keeps every auxiliary function inside the specular class. The Bochner identity used in the whole-space argument is replaced by a weighted Reilly formula, whose boundary contribution involves the second fundamental form of $\partialΩ$ and is nonnegative for convex $Ω$. Finally, we extend our results to the setting where the domain $Ω$ is non-convex. We obtain an explicit contraction rate that depends on the domain.

math.PR

Moderate Deviations for Nonlinear Hawkes Processes

A Hawkes process is a simple point process whose intensity depends on its history; the resulting dynamics are generally non-Markovian. We establish a sample-path moderate deviation principle for a nonlinear Hawkes process in the full moderate regime. Since a Poisson cluster representation is unavailable for nonlinear Hawkes processes, we use the past configuration as a Markov state and construct a potential, or Poisson corrector, for the centered stochastic intensity. A monotone Poisson coupling shows that the add-one increment of the corrector is uniformly bounded. The centered counting process is consequently the sum of a martingale with bounded jumps and an exponentially negligible boundary term. Exponential stabilization of the predictable quadratic variation follows from the process-level large deviation principle for nonlinear Hawkes processes. The martingale moderate deviation theorem then yields the result for every scale between the central-limit and large-deviation scales. The same construction gives a response formula for the asymptotic variance and, in particular, verifies that the variance dominates the stationary mean intensity in the self-exciting case.

math.PR

BRIDLE: Generalized Self-supervised Learning with Quantization

Self-supervised learning has been a powerful approach for learning meaningful representations from unlabeled data across various domains, reducing the reliance on large labeled datasets. Inspired by BERT's success in capturing deep bidirectional contexts in natural language processing, similar frameworks have been adapted to other modalities such as audio, with models like BEATs extending the bidirectional training paradigm to audio signals using vector quantization (VQ). However, these frameworks face challenges, notably their dependence on a single codebook for quantization, which may not capture the complex, multifaceted nature of signals. In addition, inefficiencies in codebook utilization lead to underutilized code vectors. To address these limitations, we introduce BRIDLE (Bidirectional Residual Quantization Interleaved Discrete Learning Encoder), a self-supervised encoder pretraining framework that incorporates residual quantization (RQ) into the bidirectional training process, and is generalized for pretraining with audio, image, and video. Using multiple hierarchical codebooks, RQ enables fine-grained discretization in the latent space, enhancing representation quality. BRIDLE involves an interleaved training procedure between the encoder and tokenizer. We evaluate BRIDLE on audio understanding tasks using classification benchmarks, achieving state-of-the-art results, and demonstrate competitive performance on image classification and video classification tasks, showing consistent improvements over traditional VQ methods in downstream performance.

cs.LG

Improved Analysis for Hessian-free High-resolution Monte Carlo Sampling

Hessian-free high-resolution (HFHR) dynamics augments underdamped Langevin dynamics (ULD) with reversible position diffusion for sampling problems that arise in machine learning. We establish an explicit quantitative contraction rate for HFHR dynamics under a position Poincaré inequality, weighted Hessian and Laplacian bounds, and a compact Sobolev embedding, where the potential function is not necessarily convex. An adapted time-augmented Poincaré inequality yields an explicit rate that improves upon the contraction rate of the underdamped Langevin dynamics. We also give a weak-solution construction and a self-contained spectral proof of the divergence lemma underlying the argument. For HFHR Monte Carlo (HFHRMC) algorithm, which is based on a discretization scheme of HFHR dynamics, we use a path-space Girsanov argument to obtain a non-asymptotic convergence bound and an explicit iteration complexity in total variation distance. The bounds hold for every $α\geq0$ and $γ>0$ and remain regular at the ULD endpoint. Optimizing the iteration complexity bound yields a positive, accuracy-dependent position-diffusion parameter at finite accuracy, while its leading high-accuracy order coincides with that of the optimized ULD endpoint. Our iteration complexity bound improves upon the existing work on HFHR algorithms. Numerical experiments including Bayesian learning problems on real data are provided to illustrate the effect of positive $α$ and its benefit.

stat.ML

DIGing--SGLD: Decentralized and Scalable Langevin Sampling over Time--Varying Networks

Sampling from a target distribution induced by training data is central to Bayesian learning, with Stochastic Gradient Langevin Dynamics (SGLD) serving as a key tool for scalable posterior sampling and decentralized variants enabling learning when data are distributed across a network of agents. This paper introduces DIGing-SGLD, a decentralized SGLD algorithm designed for scalable Bayesian learning in multi-agent systems operating over time-varying networks. Existing decentralized SGLD methods are restricted to static network topologies, and many exhibit steady-state sampling bias caused by network effects, even when full batches are used. DIGing-SGLD overcomes these limitations by integrating Langevin-based sampling with the gradient-tracking mechanism of the DIGing algorithm, originally developed for decentralized optimization over time-varying networks, thereby enabling efficient and bias-free sampling without a central coordinator. To our knowledge, we provide the first finite-time non-asymptotic Wasserstein convergence guarantees for decentralized SGLD-based sampling over time-varying networks, with explicit constants. Under standard strong convexity and smoothness assumptions, DIGing-SGLD achieves geometric convergence to an $O(\sqrtη)$ neighborhood of the target distribution, where $η$ is the stepsize, with dependence on the target accuracy matching the best-known rates for centralized and static-network SGLD algorithms using constant stepsize. Numerical experiments on Bayesian linear and logistic regression validate the theoretical results and demonstrate the strong empirical performance of DIGing-SGLD under dynamically evolving network conditions.

math.OC

An Eyring--Kramers Law for the Hypoelliptic Third-Order Langevin Diffusion

We prove an Eyring--Kramers law for metastable transitions of the hypoelliptic third-order Langevin diffusion in the low-temperature limit. This diffusion is a three-level Markovian lifting of Langevin dynamics: the Brownian noise acts only on the highest auxiliary variable and reaches the position variable through a third-order H"ormander chain. For a double-well potential with a unique index-one transition saddle, we determine both the Arrhenius exponential scale and the sharp prefactor of the mean transition time. The prefactor is governed by the unique positive unstable rate of the deterministic linearization at the saddle, equivalently the positive root of a cubic polynomial. Our proof combines a weak-capacity framework with a saddle-adapted boundary layer, an explicit Gaussian current calculation, committor localization, and intrawell flatness. Under matched kinetic normalizations, the resulting metastable prefactor is strictly smaller than its underdamped counterpart. A numerical experiment for a one-dimensional double well illustrates the Arrhenius scaling and the predicted prefactor comparison.

math.PR

Microstructural Foundation for the Rough Hawkes--Heston Model

Hawkes-based microstructural foundations for rough volatility, leverage, and rough Heston-type limits were developed by El Euch et al. (2018, Finance Stoch., 22(2), 241--280) and connected to the affine rough Heston framework of El Euch and Rosenbaum (2019, Math. Finance, 29(1), 3--38). The rough Hawkes--Heston model with common price--volatility jumps of Bondi et al. (2024, Math. Finance, 34(4), 1197--1241) extends this framework by adding state-dependent common jumps to rough affine volatility. We provide a microstructural foundation for its variance and common-jump mechanism by constructing a Poisson-embedded marked Hawkes order-flow model. Ordinary arrivals generate rough continuous volatility and leverage through a nearly unstable heavy-tailed Hawkes mechanism, while rare marked arrivals represent common shock events that produce simultaneous price jumps and volatility excitation. Under the nearly unstable scaling and the reduced-form admissibility conditions, the complete rescaled price/variance/jump system converges along the full sequence to the unique complete canonical rough Hawkes--Heston weak solution. The Hawkes renewal structure yields a Mittag--Leffler Volterra representation, which is then rewritten in Riemann--Liouville fractional form. The limiting coefficients are expressed explicitly in terms of the microscopic parameters. The construction provides a microstructural foundation for the variance and common-jump mechanism of the rough Hawkes--Heston model. Numerical experiments illustrate the convergence of our microstructural foundation to the rough Hawkes-Heston model.

q-fin.MF

Weak Equilibrium Measures and Capacity--Hitting Identities for the Hypoelliptic Third-Order Langevin Diffusion

We construct weak equilibrium measures and weak capacities for the hypoelliptic third-order Langevin diffusion motivated by an accelerated sampling algorithm (Mou et al. (2021) \textit{J. Mach. Learn. Res.}, \textbf{22}(42), 1--41). In this process, the Brownian noise acts only in the highest-order auxiliary variable and reaches the physical variables through a step-three Hörmander chain, so the standard uniformly elliptic boundary-flux theory is not directly applicable at characteristic points of phase-space balls. We prove an elliptic-regularization stability theorem for the corresponding hitting laws and then define the weak equilibrium measure and weak capacity. The proof combines the boundary-hitting stability strategy of Lee--Ramil--Seo (2026, \textit{arXiv:2503.12610v2}) with localized hypoelliptic heat-kernel estimates (Pigato (2022) \textit{Stoch. Process. Appl.}, \textbf{145}, 117--142) adapted to the third-order chain. We obtain the bounded-domain weak capacity--hitting identity and a Lyapunov drift argument in the spirit of Lee--Ramil--Seo that yields positive Harris recurrence and extends the construction to a whole-space weak equilibrium measure, and whole-space capacity--hitting identity.

math.PR

Variance Reduction for Stochastic Gradient Generalized Non-reversible Langevin Monte Carlo Algorithms

We study the leading-order fluctuation of stochastic gradient Euler-Maruyama estimators for generalized non-reversible Langevin dynamics. Under structural assumptions tailored to the small-stepsize central limit theorem and under an unbiased stochastic gradient oracle, we prove that the empirical average over a horizon of order the inverse squared stepsize satisfies a central limit theorem in the vanishing-stepsize regime. The limiting variance is characterized through the Poisson equation of the limiting full-gradient diffusion. We then rewrite this constant in an operator form that links it to the continuous-time asymptotic variance and, under standard operator-theoretic assumptions, derive a sufficient condition under which an anti-symmetric perturbation strictly reduces the leading-order fluctuation constant relative to the reversible baseline. We also identify bounded smooth predictive observables that re directly covered by the main theorem. As a separate Gaussian calculation beyond the bounded-test-function regime, we obtain closed-form formulas for quadratic Hamiltonians and linear observables. The framework covers non-reversible Langevin dynamics and augmented-state examples including Hessian-free high-resolution dynamics and a positive-definite subclass of gradient-adjusted underdamped Langevin dynamics that allow stochastic gradients. Numerical experiments on basic examples and Bayesian linear regression using synthetic data, and Bayesian logistic regression using real data support the predicted Gaussian fluctuations and show that the non-reversible schemes consistently reduce the root mean squared error (RMSE) relative to their reversible baselines.

stat.ML

VIX options in Bergomi models

We present a study of the leading-order asymptotics for VIX option prices in Bergomi models in the short-maturity and small volatility-of-volatility regimes. Both out-of-the-money (OTM) and at-the-money (ATM) asymptotics are considered for one-factor, two-factor Bergomi and $N$-factor models. The leading-order asymptotics are obtained in closed-form, which are translated into predictions for the small-maturity asymptotics of the VIX implied volatility. Numerical illustrations are provided to illustrate the efficiency of the closed-form asymptotic formulas.

q-fin.PR

Large deviations for the mean-field limit of Hawkes processes

Hawkes processes are a class of simple point processes whose intensity depends on the past history, and is in general non-Markovian. Limit theorems for Hawkes processes in various asymptotic regimes have been studied in the literature. In this paper, we study a multidimensional nonlinear Hawkes process in the asymptotic regime when the dimension goes to infinity, whose mean-field limit is a time-inhomogeneous Poisson process, and our main result is a large deviation principle for the mean-field limit.

math.PR

Accelerating Langevin Monte Carlo Sampling: A Large Deviations Analysis

Langevin algorithms are popular Markov chain Monte Carlo methods that are often used to solve high-dimensional large-scale sampling problems in machine learning. The most classical Langevin Monte Carlo algorithm is based on the overdamped Langevin dynamics. There are many variants of Langevin dynamics that often show superior performance in practice. In this paper, we provide a unified approach to study the acceleration of the variants of the overdamped Langevin dynamics through the lens of large deviations theory. Numerical experiments using both synthetic and real data are provided to illustrate the efficiency of these variants.

math.PR

Stochastic Transition-Map Distillation for Fast Probabilistic Inference

Diffusion models achieve strong generation quality, diversity, and distribution coverage, but their performance often comes with expensive inference. In this work, we propose Stochastic Transition-Map Distillation (STMD), a teacher-free framework for accelerating diffusion model inference while preserving probabilistic sample generation. In contrast to score-based diffusion models, whose denoising parametrization models the mean of the posterior distribution, STMD distills the full transition map associated with the sampling stochastic differential equation (SDE). We parameterize these SDE transitions with a conditional Mean Flow model, yielding a one- or few-step stochastic sampler that retains the transition structure of the underlying diffusion process. This perspective is especially useful for downstream tasks that require stochastic inference, such as diffusion posterior sampling, inverse problems, and energy-based fine-tuning. Compared to recent distillation methods, STMD requires no pretrained teacher, bi-level optimization, or trajectory simulation and caching, enabling efficient and scalable training. We derive convergence bounds for our method in the Wasserstein distance, providing a strong theoretical foundation for our approach, and validate STMD on various image generation examples on the MNIST, CIFAR-10, and CelebA datasets.

cs.LG

Decentralized Proximal Stochastic Gradient Langevin Dynamics

We propose Decentralized Proximal Stochastic Gradient Langevin Dynamics (DE-PSGLD), a decentralized Markov chain Monte Carlo (MCMC) algorithm for sampling from a log-concave probability distribution constrained to a convex domain. Constraints are enforced through a shared proximal regularization based on the Moreau-Yosida envelope, enabling unconstrained updates while preserving consistency with the target constrained posterior. We establish non-asymptotic convergence guarantees in the 2-Wasserstein distance for both individual agent iterates and their network averages. Our analysis shows that DE-PSGLD converges to a regularized Gibbs distribution and quantifies the bias introduced by the proximal approximation. We evaluate DE-PSGLD for different sampling problems on synthetic and real datasets. As the first decentralized approach for constrained domains, our algorithm exhibits fast posterior concentration and high predictive accuracy.

stat.ML

Rough Heston model as the scaling limit of bivariate cumulative heavy-tailed INAR processes: Weak-error bounds and option pricing

We study nearly unstable bivariate cumulative heavy-tailed INAR($\infty$) processes and show that, under a one-factor parameterization and a suitable scaling, they converge to the rough Heston model. This yields a discrete-time microstructural route to the joint price-variance dynamics and gives explicit formulas linking the INAR asymmetry parameters to the leverage correlation and diffusion scale of the limiting volatility process. On the pricing side, we derive the exact finite-$τ$ transform recursion and reduce it, in the diffusive scaling regime, to a quadratic discrete Volterra equation. We then compare this discrete equation with the continuous fractional Riccati equation from the rough Heston model. Under an admissible-strip assumption and local-in-frequency bounds, we obtain weak-error estimates for the truncated Carr--Madan pricing functional on bounded frequency windows of the form $C_1τ^{-α}+C_2(α)τ^{-(1-α)}$, where the second branch comes from the discrete-to-continuous Volterra comparison. The coefficient $C_2(α)$ collects the vanishing contributions arising from both the weakly singular baseline quadrature and the discrete-to-continuous resolvent comparison, and satisfies $C_2(α)\to0$ as $α\uparrow1^-$. We also develop an FFT-accelerated CDQ simulator with $\mathcal O(τ\log^2τ)$ complexity per path and use it to price European and path-dependent options, examine the classical limit $α=1$, and illustrate implied-volatility diagnostics.

math.PR

Accelerating Constrained Sampling: A Large Deviations Approach

The problem of sampling a target probability distribution on a constrained domain arises in many applications including machine learning. For constrained sampling, various Langevin algorithms such as projected Langevin Monte Carlo (PLMC), based on the discretization of reflected Langevin dynamics (RLD) and more generally skew-reflected non-reversible Langevin Monte Carlo (SRNLMC), based on the discretization of skew-reflected non-reversible Langevin dynamics (SRNLD), have been proposed and studied in the literature. This work focuses on the long-time behavior of SRNLD, where a skew-symmetric matrix is added to RLD. Although acceleration for SRNLD has been studied, it is not clear how one should design the skew-symmetric matrix in the dynamics to achieve good performance in practice. We establish a large deviation principle (LDP) for the empirical measure of SRNLD when the skew-symmetric matrix is chosen such that its product with the outward unit normal vector field on the boundary is zero. By explicitly characterizing the rate functions, we show that this choice of the skew-symmetric matrix accelerates the convergence to the target distribution compared to RLD and reduces the asymptotic variance. Numerical experiments for SRNLMC based on the proposed skew-symmetric matrix show superior performance, which validate the theoretical findings from the large deviations theory.

stat.ML

VIX and European options with jumps in the short-maturity regime

We present a study of the short-maturity asymptotics for VIX and European option prices in local-stochastic volatility models with compound Poisson jumps. Both out-of-the-money (OTM) and at-the-money (ATM) asymptotics are considered. The leading-order asymptotics are obtained in closed-form. We apply our results to three examples: the Eraker model, a Kou-type model, and a folded normal model. Numerical illustrations are provided for these three examples that show the accuracy of predictions based on the asymptotic results.

q-fin.PR

Sampling non-log-concave densities via Hessian-free high-resolution dynamics

We study the problem of sampling from a target distribution $π(q)\propto e^{-U(q)}$ on $\mathbb{R}^d$, where $U$ can be non-convex, via the Hessian-free high-resolution (HFHR) dynamics, which is a second-order Langevin-type process that has $e^{-U(q)-\frac12|p|^2}$ as its unique invariant distribution, and it reduces to kinetic Langevin dynamics (KLD) as the resolution parameter $α\to0$. The existing theory for HFHR dynamics in the literature is restricted to strongly-convex $U$, although numerical experiments are promising for non-convex settings as well. We focus on studying the convergence of HFHR dynamics when $U$ can be non-convex, which bridges a gap between theory and practice. Under a standard assumption of dissipativity and smoothness on $U$, we adopt the reflection/synchronous coupling method. This yields a Lyapunov-weighted Wasserstein distance in which the HFHR semigroup is exponentially contractive for all sufficiently small $α>0$ whenever KLD is. We further show that, under an additional assumption that asymptotically $\nabla U$ has linear growth at infinity, the contraction rate for HFHR dynamics is strictly better than that of KLD, with an explicit gain. As a case study, we verify the assumptions and the resulting acceleration for three examples: a multi-well potential, Bayesian linear regression with $L^p$ regularizer and Bayesian binary classification. We conduct numerical experiments based on these examples, as well as an additional example of Bayesian logistic regression with real data processed by the neural networks, which illustrates the efficiency of the algorithms based on HFHR dynamics and verifies the acceleration and superior performance compared to KLD.

math.PR