SearcharxivSearch

arXiv subjects

Lihu Xu

Publications and source records attributed to Lihu Xu.

At least 19 recordsLinked to original sources

Quantitative bounds for high dimensional entropic CLT

By extending the Johnson--Barron projection method from one dimension to high dimensions and utilizing a Wang type dimension-free Harnack inequality, we obtain a new quantitative bound for the entropic central limit theorem under the assumption that the Poincar\'e inequality holds. We compare our results with recent developments to demonstrate the merits of our approach.

math.PR

Berry-Esseen Bounds and Moderate Deviations for Catoni-Type Robust Estimation

A powerful robust mean estimator introduced by Catoni (2012) allows for mean estimation of heavy-tailed data while achieving the performance characteristics of classical mean estimator for sub-Gaussian data. While Catoni's framework has been widely extended across statistics, stochastic algorithms, and machine learning, fundamental asymptotic questions regarding the Central Limit Theorem and rare event deviations remain largely unaddressed. In this paper, we investigate Catoni-type robust estimators in two contexts: (i) mean estimation for heavy-tailed data, and (ii) linear regression with heavy-tailed innovations. For the first model, we establish the Berry--Esseen bound and moderate deviation principles, addressing both known and unknown variance settings. For the second model, we demonstrate that the associated estimator is consistent and satisfies a multi-dimensional Berry-Esseen bound.

math.ST

Long term convergence rate of Smoluchowski-Kramers approximation by Stein's method

We consider the following second-order stochastic differential equation on $\mathbb{R}^{2d}$: \begin{equation*} dX_t^m=Y_t^mdt, \quad mdY_t^m=b(X_t^m)dt+\sigma(X_t^m)dB_t-Y^m_tdt, \end{equation*} where $X^m_t$ and $Y^m_t$ represent the position and velocity of a particle at time $t$, $m>0$ denotes its mass, $b:\mathbb{R}^d \rightarrow \mathbb{R}^d$ is the drift field, $\sigma:\mathbb{R}^d \rightarrow \mathbb{R}^{d \times d}$ is the diffusion coefficient, and $\{B_t\}_{t \ge 0}$ is a $d$-dimensional standard Brownian motion. The Smoluchowski--Kramers approximation states that as $m \rightarrow 0$, this system converges to the limiting equation: \begin{equation*} dX_t=b(X_t)dt+\sigma(X_t)dB_t. \end{equation*} Utilizing Stein's method, we prove that the $1$-Wasserstein distance between the invariant distribution of $X_t^m$ and that of its small-mass limit $X_t$ is of order $O(\sqrt{m}|\ln m|).$ Particularly, in the one-dimensional case, the convergence rate can be improved to $O(\sqrt{m}).$

math.PR

Tuning free Catoni type joint robust estimation

This paper develops a Catoni-type joint (tuning-free) estimation framework for parametric models with heavy-tailed noise, in which the target parameter and the unknown noise variance are estimated simultaneously through a system of two coupled Catoni-type estimating equations. We instantiate the framework in three canonical settings: mean estimation, linear regression, and $\ell_{2}$-penalized regression. Theoretically, we establish non-asymptotic, sub-Gaussian-type deviation bounds that hold jointly for the target parameter and the variance estimator, under only a finite $2\beta$-th moment assumption with $\beta\in (1,2]$. The resulting rates match -- up to absolute constants -- those of oracle procedures that know the variance in advance, thereby attaining optimality in the heavy-tailed regime. Methodologically, because the coupled equations are intrinsically non-convex and non-linear, classical convex M-estimation arguments are inapplicable. We develop a new analytical toolkit based on the Poincare--Miranda theorem. The resulting proof strategy is of independent methodological interest, and we expect it to be applicable to a broad class of other statistical problems in which several parameters of heterogeneous nature must be estimated jointly.

math.ST

Functional central limit theorem for Euler--Maruyama scheme with decreasing step sizes

We consider the Euler--Maruyama (EM) scheme of a family of dissipative SDEs, whose step sizes $\eta_{1}\ge\eta_{2}\ge \cdots$ are decreasing, and prove that the EM scheme weakly converges to a subordinated Brownian motion $\{B_{a(t)}\}_{0\le t\le 1}$ rather than $\{B_{t}\}_{0\le t\le 1}$, where $a(t)$ is an increasing function depending on $\{\eta_{k}\}_{k \ge 1}$, for instance, $a(t)=t^{1+\alpha}$ if $\eta_k =k^{-\alpha}$. Compared to the EM scheme with constant step size, there are substantial differences as the following: (i) the EM time series is inhomogeneous and weakly converges to the ergodic measure in a polynomial speed; (ii) we have a special number $T_n =\frac{1}{\eta_1 }+\cdots+\frac{1}{\eta_n }$ which roughly measures the dependence of the EM time series; (iii) the normalized number in the CLT is $T_n ^{-1/2}n$ rather than $\sqrt{n}$, in particular, $T_n ^{-1/2}n \propto n^{(1-\beta)/2}$ when $\eta_{k}=1/k^{\beta}$ with $\beta\in(0,1)$; (iv) in the critical choice $\eta_{k}=1/k$, we have $T_{n}^{-1/2}n=O(1)$ and thus conjecture that the CLT and FCLT do not hold. This conjecture has been verified by simulations. A key distinction arises between the constant and decreasing step size implementations of the EM scheme. Under a constant step size, the time series is homogeneous. This allows one to use a stationary initialization, which automatically eliminates several complex terms in the subsequent proof of the CLT. Conversely, the time series generated by the EM scheme with decreasing step sizes forms an inhomogeneous Markov chain. To manage the analogous difficult terms in this case, that is, when the test function $h$ is Lipschitz, we must instead establish a bound for the Wasserstein-2 distance $W_{2}(\theta_k ,X_{t_k })$. This technique for handling the inhomogeneous case could be of independent interest beyond the current proof.

math.PR

$W_{\bf d}$-convergence rate of EM schemes for invariant measures of supercritical stable SDEs

By establishing the regularity estimates for nonlocal Stein/Poisson equations under $\gamma$-order H\"older and dissipative conditions on the coefficients, we derive the $W_{\bf d}$-convergence rate for the Euler-Maruyama schemes applied to the invariant measure of SDEs driven by multiplicative $\alpha$-stable noises with $\alpha \in (\frac{1}{2}, 2)$, where $W_{\bf d}$ denotes the Wasserstein metric with ${\bf d}(x,y)=|x-y|^\gamma\wedge 1$ and $\gamma \in ((1-\alpha)_+, 1]$.

math.PR

Robust estimation for high-dimensional time series with heavy tails

We study in this paper the problem of least absolute deviation (LAD) regression for high-dimensional heavy-tailed time series which have finite $\alpha$-th moment with $\alpha \in (1,2]$. To handle the heavy-tailed dependent data, we propose a Catoni type truncated minimization problem framework and obtain an $\mathcal{O}\big( \big( (d_1+d_2) (d_1\land d_2) \log^2 n / n \big)^{(\alpha - 1)/\alpha} \big)$ order excess risk, where $d_1$ and $d_2$ are the dimensionality and $n$ is the number of samples. We apply our result to study the LAD regression on high-dimensional heavy-tailed vector autoregressive (VAR) process. Simulations for the VAR($p$) model show that our new estimator with truncation are essential because the risk of the classical LAD has a tendency to blow up. We further apply our estimation to the real data and find that ours fits the data better than the classical LAD.

math.ST

Error estimates between SGD with momentum and underdamped Langevin diffusion

Stochastic gradient descent with momentum is a popular variant of stochastic gradient descent, which has recently been reported to have a close relationship with the underdamped Langevin diffusion. In this paper, we establish a quantitative error estimate between them in the 1-Wasserstein and total variation distances.

stat.ML

Total variation distance between SDEs with stable noise and Brownian motion

We consider a $d$-dimensional stochastic differential equation (SDE) of the form $d U_t = b(U_t) dt + \sigma\,d Z_t$, let $X_t$ be the solution if the driving noise $Z_t$ is a $d$-dimensional rotationally symmetric $\alpha$-stable process ($1<\alpha<2$), and let $Y_t$ be the solution if the driving noise is a $d$-dimensional Brownian motion. Continuing the work of [Deng,Schilling, Xu, Bernoulli, 23], we derive an estimate of the total variation distance $\|{\rm L} (X_{t})-{\rm law}(Y_{t})\|_{\rm TV}$ for all $t>0$, and we show that the ergodic measures $\mu_\alpha$ and $\mu_2$ of $X_t$ and $Y_t$, respectively, satisfy $$\|\mu_\alpha-\mu_2\|_{\rm TV} \leq \frac{Cd\log(1+d)}{\alpha-1}(2-\alpha).$$ We shall show that this bound is optimal with respect to $\alpha$ by an Ornstein--Uhlenbeck SDE. Combining this bound with a recent interpolation result from \cite{HRW23}, we can derive a bound in Wasserstein-$p$ distance ($0< p <1$): \begin{gather*} \|\mu_\alpha-\mu_2\|_{W_p} \leq\frac{Cd^{(p+3)/2}\log(1+d)}{\alpha-1} (2-\alpha). \end{gather*} {\bf Key Words:} Total variation distance, Wasserstein-$p$ distance, stochastic differential equation, Poisson equation, stable process.

math.PR

Quantitative diffusion approximation for the Neutral $r$-Alleles Wright-Fisher Model with Mutations

We apply a Lindeberg principle under the Markov process setting to approximate the Wright-Fisher model with neutral $r$-alleles using a diffusion process, deriving an error rate based on a function class distance involving fourth-order bounded differentiable functions. This error rate consists of a linear combination of the maximum mutation rate and the reciprocal of the population size. Our result improves the error bound in the seminal work [PNAS,1977], where only the special case $r=2$ was studied.

math.PR

Unbiased approximation of the ergodic measure for piecewise $\alpha$-stable Ornstein-Uhlenbeck processes arising in queueing networks

Piecewise $\alpha$-stable Ornstein-Uhlenbeck (OU) processes arising in queue networks usually do not have an explicit dissipation, which makes the related numerical methods such as Euler-Maruyama (EM) scheme more difficult to analyze. We develop an EM scheme with decreasing step size $\Lambda=(\eta_n)_{n\in \mathbb{N}}$ to approximate their ergodic measures. This approximation does not have a bias and has a rate $\eta^{1/\alpha}_n$ in Wasserstein-1 distance. We show by the classical OU process that our convergence rate is optimal. We further prove the central limit theorem (CLT) and moderate derivation principle (MDP) for the empirical measure of these piecewise $\alpha$-stable Ornstein-Uhlenbeck processes. In addition, we use the Sinkhorn--Knopp algorithm to compute the Wasserstein-1 distance and conduct simulations for several concrete examples.

math.PR

Phase transition in the EM scheme of an SDE driven by $α$-stable noises with $α\in (0,2]$

We study in this paper the EM scheme for a family of well-posed critical SDEs with the drift $-x\log(1+|x|)$ and $α$-stable noises. Specifically, we find that when the SDE is driven by a rotationally symmetric $α$-stable processes with $α=2$ (i.e. Brownian motion), the EM scheme is bounded in the $L^2$ sense uniformly w.r.t. the time. In contrast, if the SDE is driven by a rotationally symmetric $α$-stable process with $α\in (0,2)$, all the $β$-th moments, with $β\in (0,α)$, of the EM scheme blow up. This demonstrates a phase transition phenomenon as $α\uparrow 2$. We verify our results by simulations.

math.PR

Optimal Wasserstein-$1$ distance between SDEs driven by Brownian motion and stable processes

We are interested in the following two $\mathbb{R}^d$-valued stochastic differential equations (SDEs): \begin{gather*} d X_t=b(X_t)\,d t + σ\,d L_t, \quad X_0=x, %\label{BM-SDE} d Y_t=b(Y_t)\,d t + σ\,d B_t, \quad Y_0=y, \end{gather*} where $σ$ is an invertible $d\times d$ matrix, $L_t$ is a rotationally symmetric $α$-stable Lévy process, and $B_t$ is a $d$-dimensional standard Brownian motion (note that $B_t$ is a rotationally symmetric $α$-stable Lévy process with $α=2$). We show that for any $α_0 \in (1,2)$ the Wasserstein-$1$ distance $W_1$ satisfies for $α\in [α_0,2)$ \begin{gather*} W_{1}\left(X_{t}^x, Y_{t}^y\right) \leq C_1 e^{-C_2t}|x-y| +\frac{C}{α_0-1}(2-α)d\log(1+d), \end{gather*} which implies, in particular, \begin{equation} \label{e:W1Rate} W_1(μ_α, μ_2) \leq \frac{C}{α_0-1}(2-α)d\log(1+d), \end{equation} where $μ_α$ and $μ_2$ are the ergodic measures of $X_t$ and $Y_t$ respectively. For the special case of a $d$-dimensional Ornstein--Uhlenbeck system, we show that $W_1(μ_α, μ_2) \geq C_{d} (2-α)$ for all $α\in(1,2)$; this indicates that the convergence rate with respect to $α$ in the second bound is optimal. The term $d\log(1+d)$ appearing in this bound seems to be optimal for the dimension $d$ as well.

math.PR

Stable central limit theorem in total variation distance

Under certain general conditions, we prove that the stable central limit theorem holds in the total variation distance and get its optimal convergence rate for all $α\in (0,2)$. Our method is by two measure decompositions, one step estimates, and a very delicate induction with respect to $α$. One measure decomposition is light tailed and borrowed from \cite{BC16}, while the other one is heavy tailed and indispensable for lifting convergence rate for small $α$. The proof is elementary and composed of the ingredients at the postgraduate level. Our result clarifies that when $α=1$ and $X$ has a symmetric Pareto distribution, the optimal rate is $n^{-1}$ rather than $n^{-1} (\ln n)^2$ as conjectured in literatures.

math.PR

Approximation of the invariant measure for stable SDE by the Euler-Maruyama scheme with decreasing step-sizes

Let $(X_t)_{t \ge 0}$ be the solution of the stochastic differential equation $$dX_t = b(X_t) dt+A dZ_t, \quad X_{0}=x,$$ where $b: \mathbb{R}^d \rightarrow \mathbb R^d$ is a Lipschitz function, $A \in \mathbb R^{d \times d}$ is a positive definite matrix, $(Z_t)_{t\geq 0}$ is a $d$-dimensional rotationally invariant $α$-stable Lévy process with $α\in (1,2)$ and $x\in\mathbb{R}^{d}$. We use two Euler-Maruyama schemes with decreasing step sizes $Γ= (γ_n)_{n\in \mathbb{N}}$ to approximate the invariant measure of $(X_t)_{t \ge 0}$: one with i.i.d. $α$-stable distributed random variables as its innovations and the other with i.i.d. Pareto distributed random variables as its innovations. We study the convergence rate of these two approximation schemes in the Wasserstein-1 distance. For the first scheme, when the function $b$ is Lipschitz and satisfies a certain dissipation condition, we show that the convergence rate is $γ^{1/α}_n$. Under an additional assumption on the second order directional derivatives of $b$, this convergence rate can be improved to $γ^{1+\frac 1 α-\frac{1}κ}_n$ for any $κ\in [1,α)$. For the second scheme, when the function $b$ is twice continuously differentiable, we obtain a convergence rate of $γ^{\frac{2-α}α}_n$. We show that the rate $γ^{\frac{2-α}α}_n$ is optimal for the one dimensional stable Ornstein-Uhlenbeck process. Our theorems indicate that the recent remarkable result about the unadjusted Langevin algorithm with additive innovations can be extended to the SDEs driven by an $α$-stable Lévy process and the corresponding convergence rate has a similar behaviour. Compared with the previous result, we have relaxed the second order differentiability condition to the Lipschitz condition for the first scheme.

math.PR

Unadjusted Langevin Algorithms for SDEs with Hoelder Drift

Consider the following stochastic differential equation for $(X_t)_{t\ge 0}$ on $\mathbb R^d$ and its Euler-Maruyama (EM) approximation $(Y_{t_n})_{n\in \mathbb Z^+}$: \begin{align*} &d X_t=b( X_t) d t+σ(X_t) d B_t, \\ & Y_{t_{n+1}}=Y_{t_{n}}+η_{n+1} b(Y_{t_{n}})+σ(Y_{t_{n}})\left(B_{t_{n+1}}-B_{t_{n}}\right), \end{align*} where $b:\mathbb{R}^d \rightarrow \mathbb{R}^d,\ \ σ: \mathbb R^d \rightarrow \mathbb{R}^{d \times d}$ are measurable, $B_t$ is the $d$-dimensional Brownian motion, $t_0:=0,t_{n}:=\sum_{k=1}^{n} η_{k}$ for constants $η_k>0$ satisfying $\lim_{k \rightarrow \infty} η_k=0$ and $\sum_{k=1}^\inftyη_k =\infty$. Under (partial) dissipation conditions ensuring the ergodicity, we obtain explicit convergence rates of $\mathbb W_p(\mathscr{L}(Y_{t_n}), \mathscr{L}(X_{t_n}))+\mathbb W_p(\mathscr{L}(Y_{t_n}), μ)\rightarrow 0$ as $n\rightarrow \infty$, where $\mathbb W_p$ is the $L^p$-Wasserstein distance for certain $p\in [0,\infty)$, $\mathscr{L}(ξ)$ is the distribution of random variable $ξ$, and $μ$ is the unique invariant probability measure of $(X_t)_{t \ge 0}$. Comparing with the existing results where $b$ is at least $C^2$-smooth, our estimates apply to Hoelder continuous drift and can be sharp in several specific situations.

math.PR

Approximation of the invariant measure of stable SDEs by an Euler--Maruyama scheme

We propose two Euler-Maruyama (EM) type numerical schemes in order to approximate the invariant measure of a stochastic differential equation (SDE) driven by an $α$-stable Lévy process ($1<α<2$): an approximation scheme with the $α$-stable distributed noise and a further scheme with Pareto-distributed noise. Using a discrete version of Duhamel's principle and Bismut's formula in Malliavin calculus, we prove that the error bounds in Wasserstein-$1$ distance are in the order of $η^{1-ε}$ and $η^{\frac2α-1}$, respectively, where $ε\in (0,1)$ is arbitrary and $η$ is the step size of the approximation schemes. For the Pareto-driven scheme, an explicit calculation for Ornstein--Uhlenbeck $α$-stable process shows that the rate $η^{\frac2α-1}$ cannot be improved.

math.PR

Cramér-type moderate deviations for Euler-Maruyama scheme for SDE

In this paper, we establish normalized and self-normalized Cramér-type moderate deviations for Euler-Maruyama scheme for SDE. As a consequence of our results, Berry-Esseen's bounds and moderate deviation principles are also obtained. Our normalized Cramér-type moderate deviations refines the recent work of [Lu, J., Tan, Y., Xu, L., 2022. Central limit theorem and self-normalized Cramér-type moderate deviation for Euler-Maruyama scheme. Bernoulli 28(2): 937--964].

math.PR