SearcharxivSearch

arXiv subjects

Zhonggen Su

Publications and source records attributed to Zhonggen Su.

At least 19 recordsLinked to original sources

Non-symmetric vector dyson equations

We study the vector Dyson equation $$-\frac{1}{m(z)}=z\mathbf{1}+\mathbf{a}+Sm(z),$$ with parameter $z$ in the complex upper half-plane $\mathbb{C}_+$, where $\mathbf{a}\in\mathbb R^d$ and $S$ is a nonnegative matrix, not necessarily symmetric. This equation has a unique vector solution $m(z)\in\mathbb{C}_+^d$, for which we establish a complete measure decomposition and prove regularity. We then develop a graph-theoretic approach to the singularity and stability problem for non-symmetric matrices $S$. The graph structure of $S$ identifies the possible degeneracies of the stability operator as $z$ approaches the real axis. In particular, for non-backtracking matrices, we prove square-root growth at regular edges, cubic-root growth at regular cusps, and complete stability estimates. We also obtain the corresponding estimates for symmetric matrices in the periodic setting.

math.CV

Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning

Referring expression understanding in remote sensing poses unique challenges, as it requires reasoning over complex object-context relationships. While supervised fine-tuning (SFT) on multimodal large language models achieves strong performance with massive labeled datasets, they struggle in data-scarce scenarios, leading to poor generalization. To address this limitation, we propose Geo-R1, a reasoning-centric reinforcement fine-tuning (RFT) paradigm for few-shot geospatial referring. Geo-R1 enforces the model to first generate explicit, interpretable reasoning chains that decompose referring expressions, and then leverage these rationales to localize target objects. This "reason first, then act" process enables the model to make more effective use of limited annotations, enhances generalization, and provides interpretability. We validate Geo-R1 on three carefully designed few-shot geospatial referring benchmarks, where our model consistently and substantially outperforms SFT baselines. It also demonstrates strong cross-dataset generalization, highlighting its robustness. Code and data will be released at: https://github.com/Geo-R1/geo-r1.

cs.CV

Normal approximation for the polynomial functionals of correlated random field sampling along random walk path in dimension $1+1$

Let $ξ$ be the stationary occupation field generated by a Poisson system of independent simple symmetric random walks on $\mathbb Z$ in space--time dimension $1+1$. For a finite set $A\subset\mathbb Z$, we consider the classical fixed-region observables $W_N(A)$, the cumulative occupation of $A$ up to time $N$, and $D_N(A)$, the number of distinct particles visiting $A$ up to time $N$. We prove quantitative central limit theorems for both observables, with Wasserstein rate of order $N^{-1/4}$. In addition, we introduce an independent nearest-neighbour random walk $S=(S_n,\,n\ge 0)$ on $\mathbb Z$ with non-zero drift and sample the field along this ballistic path. For a fixed polynomial observable $φ(x)=\sum_{j=0}^k β_j x^j, β_k\neq 0$, of degree $k\in \mathbb N$, we consider the partial sums $Y_{N,φ}=\sum_{n=1}^N φ(ξ(n,S_n)).$ We prove a Wasserstein bound of order $N^{-1/2}$ for the normal approximation of the standardized $Y_{N,φ}$. To the best of our knowledge, this is the first quantitative normal approximation result for polynomial functionals of the Poisson occupation field sampled along a random walk path. The drift induces an effective decorrelation of the sampled environment, leading to a substantial improvement over fixed-region sampling. The proofs rely on a representation of $ξ$ as a Poisson functional on path space and on the Malliavin--Stein method for Poisson functionals.

math.PR

High-dimensional normal approximations for sums of Langevin Markov chains

Consider the well-known Langevin diffusion on $\mathbb{R}^d$ $$\mathrm{d} X_t = -\nabla U(X_t)\,\mathrm{d} t + \sqrt{2}\mathrm{d} B_t, $$ and its Euler-Maruyama discretization given by $$X_{k+1}=X_k-η\nabla U(X_k)+\sqrt{2η}ξ_{k+1},$$ where $η$ is the step size. Under mild conditions, the Langevin diffusion admits $π(\mathrm{d} x)\propto \exp(-U(x))\mathrm{d} x$ as its unique stationary distribution. In this paper, we mainly study the normal approximation of the normalized partial sum $$ W_n = η^{1/2} n^{-1/2} \left( \sum_{i=0}^{n-1} X_i- \int_{\mathbb{R}^d} x\,π(\mathrm{d} x) \right).$$ To the best of our knowledge, this work provides the first dimension-explicit convergence rates in high-dimensional settings. Our main tool is a novel upper bound for the 1-Wasserstein distance $W_1(W,γ)$ via the exchange pair approach, where $W$ is any random vector of interest and $γ$ is a $d$-dimensional standard normal random vector.

math.PR

Non-asymptotic Analysis of Poisson randomized midpoint Langevin Monte Carlo

The task of sampling from a high-dimensional distribution $π$ on $\R^d$ is a fundamental algorithmic problem with applications throughout statistics, engineering, and the sciences. Consider the Langevin diffusion on $\R^d$ \begin{align*} \dif X_t=-\nabla U(X_t)dt+\sqrt{2}dB_t, \end{align*} under mild conditions, it admits $π(\dif x)\propto \exp(-U(x))\dif x$ as its unique stationary distribution. Recently, Kandasamy and Nagaraj (2024) introduced a stochastic algorithm called Poisson Randomized Midpoint Langevin Monte Carlo (PRLMC) to enhance the rate of convergence towards the target distribution $π$. In this paper, we first show that under mild conditions, the PRLMC, as a Markov chain, admits a unique stationary distribution $π_η$ ($η$ is the step size) and obtain the convergence rate of PRLMC to $π_η$ in total variation distance. Then we establish a sharp error bound between $π_η$ and $π$ under the 2-Wasserstein distance. Finally, we propose a decreasing-step size version of PRLMC and provide its convergence rate to $π$ which is nearly optimal.

math.ST

Convergence rate of randomized midpoint Langevin Monte Carlo

The randomized midpoint Langevin Monte Carlo (RLMC), introduced by Shen and Lee (2019), is a variant of classical Unadjusted Langevin Algorithm. It was shown in the literature that the RLMC is an efficient algorithm for approximating high-dimensional probability distribution $π$. In this paper, we establish the exponential ergodicity of RLMC with constant step-size. Moreover, we design a dereasing-step size RLMC and provide its convergence rate in terms of a functional class distance.

math.ST

Coupling between Brownian motion and random walks on the infinite percolation cluster

For the supercritical Bernoulli bond percolation on $\mathbb{Z}^d$ ($d \geq 2$), we give a coupling between the random walk on the infinite cluster and its limit Brownian motion, such that the maximum distance between the paths during $[0,T]$ has a mean of order $T^{\frac{1}{3}+o(1)}$. The construction of the coupling utilizes the optimal transport tool. The analysis mainly relies on local CLT and the concentration of the cluster density. This partially answers an open question posed by Biskup [Probab. Surv., 8:294-373, 2011]. As a direct application, our result recovers the law of the iterated logarithm proved by Duminil-Copin [arXiv:0809.4380], and further identifies the limit constant.

math.PR

Conditional central limit theorems for exponential random graphs

In this paper, we study the Exponential Random Graph Models (ERGMs) conditioning on the number of edges. In subcritical region of model parameters, we prove a conditional Central Limit Theorem (CLT) with explicit mean and variance for the number of two stars. This generalizes the corresponding result in the literature for the Erdős--Rényi random graph. To prove our main result, we develop a new conditional CLT via exchangeable pairs based on the ideas of Dey and Terlov. Our key technical contributions in the application to ERGMs include establishing a linearity condition for an exchangeable pair involving two star counts, a local CLT for edge counts, as well as new higher-order concentration inequalities. Our approach also works for general subgraph counts, and we give a conjectured form of their conditional CLT.

math.PR

Quantitative estimates of the spectral norm of random matrices with independent columns

This paper investigates the nonasymptotic properties of the spectral norm of some random matrices with independent columns. In particular, we consider an $m\times n$ random matrix $BA$, where $A$ is an $N\times n$ random matrix with independent mean-zero subexponential entries, and $B$ is an $m\times N$ deterministic matrix. We prove that the $L_{p}$ norm of the spectral norm of $BA$ is upper bounded by $(\sqrt{m}+\sqrt{n})p$. It is remarkable that this result is independent of the dimension $N$.

math.PR

Quantitative estimates of the singular values of random i.i.d. matrices

Let $M$ be an $n\times n$ random i.i.d. matrix. This paper studies the deviation inequality of $s_{n-k+1}(M)$, the $k$-th smallest singular value of $M$. In particular, when the entries of $M$ are subgaussian, we show that for any $γ\in (0, 1/2), \varepsilon>0$ and $\log n\le k\le c\sqrt{n}$ \begin{align} \textsf{P}\{s_{n-k+1}(M)\le \frac{\varepsilon}{\sqrt{n}} \}\le \Big( \frac{C\varepsilon}{k}\Big)^{γk^{2}}+e^{-c_{1}kn}.\nonumber \end{align} This result improves an existing result of Nguyen, which obtained a deviation inequality of $s_{n-k+1}(M)$ with $(C\varepsilon/k)^{γk^{2}}+e^{-cn}$ decay.

math.PR

Three-Parameter Approximations of Sums of Locally Dependent Random Variables via Stein's Method

Let $\{X_{i}, i\in J\}$ be a family of locally dependent non-negative integer-valued random variables with finite expectations and variances. We consider the sum $W=\sum_{i\in J}X_i$ and use Stein's method to establish general upper error bounds for the total variation distance $d_{TV}(W, M)$, where $M$ represents a three-parameter random variable. As a direct consequence, we obtain a discretized normal approximation for $W$. As applications, we study in detail four well-known examples, which are counting vertices of all edges point inward, birthday problem, counting monochromatic edges in uniformly colored graphs, and triangles in the Erdős-Rényi random graph. Through delicate analysis and computations, we obtain sharper upper error bounds than existing results.

math.PR

Deviation Inequalities for the Spectral Norm of Structured Random Matrices

We study the deviation inequality for the spectral norm of structured random matrices with non-gaussian entries. In particular, we establish an optimal bound for the $p$-th moment of the spectral norm by transfering the spectral norm into the suprema of canonical processes. A crucial ingredient of our proof is a comparison of weak and strong moments. As an application, we show a deviation inequality for the smallest singular value of a rectangular random matrix.

math.PR

On Log-Concave-Tailed Chaoses and the Restricted Isometry Property

In this paper, we obtain a $p$-th moment bound for the suprema of a log-concave-tailed nonhomogeneous chaos process, which is optimal in some special cases. A crucial ingredient of the proof is a novel decoupling inequality, which may be of independent interest. With this $p$-th moment bound, we show two uniform Hanson-Wright type deviation inequalities for $α$-subexponential entries ($1\le α\le 2$), which recover some known results. As applications, we prove the restricted isometry property of partial random circulant matrices and time-frequency structured random matrices induced by standard $α$-subexponential vectors ($1\le α\le 2$), which extends the previously known results for the subgaussian case.

math.PR

Uniform Hanson-Wright Type Deviation Inequalities for $α$-Subexponential Random Vectors

This paper is devoted to uniform versions of the Hanson-Wright inequality for a random vector with independent centered $α$-subexponential entries, $0<α\le 1$. Our method relies upon a novel decoupling inequality and a comparison of weak and strong moments. As an application, we use the derived inequality to prove the restricted isometry property of partial random circulant matrices generated by standard $α$-subexponential random vectors, $0<α\le 1$.

math.PR

Breaking the Communication-Privacy-Accuracy Tradeoff with $f$-Differential Privacy

We consider a federated data analytics problem in which a server coordinates the collaborative data analysis of multiple users with privacy concerns and limited communication capability. The commonly adopted compression schemes introduce information loss into local data while improving communication efficiency, and it remains an open problem whether such discrete-valued mechanisms provide any privacy protection. In this paper, we study the local differential privacy guarantees of discrete-valued mechanisms with finite output space through the lens of $f$-differential privacy (DP). More specifically, we advance the existing literature by deriving tight $f$-DP guarantees for a variety of discrete-valued mechanisms, including the binomial noise and the binomial mechanisms that are proposed for privacy preservation, and the sign-based methods that are proposed for data compression, in closed-form expressions. We further investigate the amplification in privacy by sparsification and propose a ternary stochastic compressor. By leveraging compression for privacy amplification, we improve the existing methods by removing the dependency of accuracy (in terms of mean square error) on communication cost in the popular use case of distributed mean estimation, therefore breaking the three-way tradeoff between privacy, communication, and accuracy. Finally, we discuss the Byzantine resilience of the proposed mechanism and its application in federated learning.

cs.CR

Approximation of Sums of Locally Dependent Random Variables via Perturbation of Stein Operator

Let $(X_{i}, i\in J)$ be a family of locally dependent nonnegative integer-valued random variables, and consider the sum $W=\sum\nolimits_{i\in J}X_i$. We first establish a general error upper bound for $d_{TV}(W, M)$ using Stein's method, where the target variable $M$ is either the mixture of Poisson distribution and binomial or negative binomial distribution. As applications, we attain $O(|J|^{-1})$ error bounds for ($k_{1},k_{2}$)-runs and $k$-runs under some special cases. Our results are significant improvements of the existing results in literature, say $O(|J|^{-0.5})$ in Peköz [Bernoulli, 19 (2013)] and $O(1)$ in Upadhye, et al. [Bernoulli, 23 (2017)].

math.PR

Controlling FSR in Selective Classification

Uncertainty quantification and false selection error rate (FSR) control are crucial in many high-consequence scenarios, so we need models with good interpretability. This article introduces the optimality function for the binary classification problem in selective classification. We prove the optimality of this function in oracle situations and provide a data-driven method under the condition of exchangeability. We demonstrate it can control global FSR with the finite sample assumption and successfully extend the above situation from binary to multi-class classification. Furthermore, we demonstrate that FSR can still be controlled without exchangeability, ultimately completing the proof using the martingale method.

math.ST

Rates of convergence in the distances of Kolmogorov and Wasserstein for standardized martingales

We give some rates of convergence in the distances of Kolmogorov and Wasserstein for standardized martingales with differences having finite variances. For the Kolmogorov distances, we present some exact Berry-Esseen bounds for martingales, which generalizes some Berry-Esseen bounds due to Bolthausen. For the Wasserstein distance, with Stein's method and Lindeberg's telescoping sum argument, the rates of convergence in martingale central limit theorems recover the classical rates for sums of i.i.d.\ random variables, and therefore they are believed to be optimal.

math.PR