SearcharxivSearch

arXiv subjects

Xiucai Ding

Publications and source records attributed to Xiucai Ding.

At least 19 recordsLinked to original sources

Phase transitions and approximations of mean squared error for state-space models with fractional differencing

We study trend estimation in state-space models in which the trend has a fractional stochastic difference of order $d>0$ and the observation errors form a short-range-dependent stationary process. Using finite-sequence fractional summation and differencing operators, we analyze the penalized least-squares estimator obtained by shrinking the fractional differences of the trend. We derive asymptotic mean squared error (MSE) approximations for all $d>0$ and identify a sharp phase transition at $d=1/2$. When $d>1/2$, the estimator is consistent and its optimally balanced MSE has order $n^{-(2d-1)/(2d)}$. At the boundary $d=1/2$, we obtain a refined finite-sample approximation and show that the MSE decreases at the slower order $\log\log n/\log n$. When $0<d<1/2$, the MSE converges to an explicit positive limit, so consistent recovery of the trend is impossible under the considered scaling. We also describe a practical criterion for choosing the penalty parameter and differencing order, and numerical experiments illustrate the MSE approximations and the behavior of the selection procedure.

math.ST

Bias-Corrected Multiplier Bootstrap Inference for Spectral Edges of Large Covariance Matrices

Inference for spectral edges of large covariance matrices is a fundamental problem in high-dimensional statistics. A major difficulty is that the largest non-spiked sample eigenvalues, which serve as natural estimators of the edge, fluctuate on the Tracy--Widom scale. Consequently, valid inference requires accurate centering by the deterministic spectral edge together with a precise scaling constant, both of which are often difficult to estimate in practice under general unknown population covariance structures. In this paper, we propose a bias-corrected multiplier bootstrap procedure for inference on the deterministic edge of the bulk spectrum. The key idea is to introduce a carefully calibrated multiplier perturbation that regularizes the edge fluctuation to a slightly larger scale at which Gaussian approximation becomes tractable. The resulting confidence interval is constructed directly from bootstrap eigenvalues, together with a data-driven recentering step that corrects the bootstrap-induced shift of the deterministic edge. On the theoretical side, we show that, after bias correction and rescaling, the largest few non-spiked bootstrap eigenvalues are asymptotically Gaussian conditionally on the data. Building on this result, we establish the asymptotic validity of the proposed confidence interval, whose length is only slightly larger than the Tracy--Widom scale, and prove vanishing coverage under alternatives in which additional spikes separate from the bulk at a local scale larger than $n^{-1/6}$. As a consequence, the same confidence interval yields a threshold-free estimator for the number of spikes, without requiring the spikes to be distinct or very large. Equivalently, the procedure yields a data-driven and theoretically justified cutoff for the scree plot.

stat.ME

A spectral inference method for determining the number of communities in networks

To characterize the community structure in network data, researchers have developed various block-type models, including the stochastic block model, the degree-corrected stochastic block model, the mixed membership block model, the degree-corrected mixed membership block model, and others. A critical step in applying these models effectively is determining the number of communities in the network. However, to the best of our knowledge, existing methods for estimating the number of network communities either rely on explicit model fitting or fail to simultaneously accommodate network sparsity and a diverging number of communities. In this paper, we propose a model-free spectral inference method based on eigengap ratios that addresses these challenges. The inference procedure is straightforward to compute, requires no parameter tuning, and can be applied to a wide range of block models without the need to estimate network distribution parameters. Furthermore, it is effective for both dense and sparse networks with a divergent number of communities. Technically, we show that the proposed spectral test statistic converges to a {function of the type-I Tracy-Widom distribution via the Airy kernel} under the null hypothesis, and that the test is asymptotically powerful under weak alternatives. Simulation studies on both dense and sparse networks demonstrate the efficacy of the proposed method. Three real-world examples are presented to illustrate the usefulness of the proposed test.

stat.ME

Generalized Robust Adaptive-Bandwidth Multi-View Manifold Learning in High Dimensions with Noise

Multiview datasets are common in scientific and engineering applications, yet existing fusion methods offer limited theoretical guarantees, particularly in the presence of heterogeneous and high-dimensional noise. We propose Generalized Robust Adaptive-Bandwidth Multiview Diffusion Maps (GRAB-MDM), a new kernel-based diffusion geometry framework for integrating multiple noisy data sources. The key innovation of GRAB-MDM is a {view}-dependent bandwidth selection strategy that adapts to the geometry and noise level of each view, enabling a stable and principled construction of multiview diffusion operators. Under a common-manifold model, we establish asymptotic convergence results and show that the adaptive bandwidths lead to provably robust recovery of the shared intrinsic structure, even when noise levels and sensor dimensions differ across views. Numerical experiments demonstrate that GRAB-MDM significantly improves robustness and embedding quality compared with fixed-bandwidth and equal-bandwidth baselines, and usually outperform existing algorithms. The proposed framework offers a practical and theoretically grounded solution for multiview sensor fusion in high-dimensional noisy environments.

stat.ML

On the edge eigenvalues of sparse random geometric graphs

In this paper, we study the edge eigenvalues of random geometric graphs (RGGs) generated by multivariate Gaussian samples in the sparse regime under a broad class of distance metrics. Previous work on edge eigenvalues under related setups has relied on methods based on integral operators or the Courant-Fischer min-max principle with interpolation. However, these approaches typically require either a dense regime or sampling distributions that are compactly supported and non-vanishing, and therefore cannot be generalized to our setting. We introduce a two-step smoothing-matching argument. First, we construct a smoothed empirical operator from the given RGG. We then match its edge eigenvalues to those of a continuum limit operator via a counting argument. We show that, after proper normalization, the first few nontrivial edge eigenvalues of the RGG converge with high probability to those of a differential operator, which can be computed explicitly through a simple second-order linear partial differential equation. To the best of our knowledge, these are the first results on the edge eigenvalues of RGGs generated from samples with unbounded support and vanishing density functions.

math.PR

Structural Classification of Locally Stationary Time Series Based on Second-order Characteristics

Time series classification is crucial for numerous scientific and engineering applications. In this article, we present a numerically efficient, practically competitive, and theoretically rigorous classification method for distinguishing between two classes of locally stationary time series based on their time-domain, second-order characteristics. Our approach builds on the autoregressive approximation for locally stationary time series, combined with an ensemble aggregation and a distance-based threshold for classification. It imposes no requirement on the training sample size, and is shown to achieve zero misclassification error rate asymptotically when the underlying time series differ only mildly in their second-order characteristics. The new method is demonstrated to outperform a variety of state-of-the-art solutions, including wavelet-based, tree-based, convolution-based methods, as well as modern deep learning methods, through intensive numerical simulations and a real EEG data analysis for epilepsy classification.

stat.ME

Simultaneous Sieve Estimation and Inference for Time-Varying Nonlinear Time Series Regression

In this paper, we investigate time-varying nonlinear time series regression for a broad class of locally stationary time series. First, we propose sieve nonparametric estimators for the time-varying regression functions that achieve uniform consistency. Second, we develop a unified simultaneous inferential theory to conduct both structural and exact form tests on these functions. Additionally, we introduce a multiplier bootstrap procedure for practical implementation. Our methodology and theory require only mild assumptions on the regression functions, allow for unbounded domain support, and effectively address the issue of identifiability for practical interpretation. Technically, we establish sieve approximation theory for 2-D functions in unbounded domains, prove two Gaussian approximation results for affine forms of high-dimensional locally stationary time series, and calculate critical values for the maxima of the Gaussian random field arising from locally stationary time series, which may be of independent interest. Numerical simulations and two data analyses support our results, and we have developed an $\mathtt{R}$ package, $\mathtt{SIMle}$, to facilitate implementation.

stat.ME

A Lanczos-Based Algorithmic Approach for Spike Detection in Large Sample Covariance Matrices

We introduce a new approach for estimating the number of spikes in a general class of spiked covariance models without directly computing the eigenvalues of the sample covariance matrix. This approach is based on the Lanczos algorithm and the asymptotic properties of the associated Jacobi matrix and its Cholesky factorization. A key aspect of the analysis is interpreting the eigenvector spectral distribution as a perturbation of its asymptotic counterpart. The specific exponential-type asymptotics of the Jacobi matrix enables an efficient approximation of the Stieltjes transform of the asymptotic spectral distribution via a finite continued fraction. As a consequence, we also obtain estimates for the density of the asymptotic distribution and the location of outliers. We provide consistency guarantees for our proposed estimators, proving their convergence in the high-dimensional regime. We demonstrate that, when applied to standard spiked covariance models, our approach outperforms existing methods in computational efficiency and runtime, while still maintaining robustness to exotic population covariances.

math.ST

Kernel spectral joint embeddings for high-dimensional noisy datasets using duo-landmark integral operators

Integrative analysis of multiple heterogeneous datasets has become standard practice in many research fields, especially in single-cell genomics and medical informatics. Existing approaches oftentimes suffer from limited power in capturing nonlinear structures, insufficient account of noisiness and effects of high-dimensionality, lack of adaptivity to signals and sample sizes imbalance, and their results are sometimes difficult to interpret. To address these limitations, we propose a novel kernel spectral method that achieves joint embeddings of two independently observed high-dimensional noisy datasets. The proposed method automatically captures and leverages possibly shared low-dimensional structures across datasets to enhance embedding quality. The obtained low-dimensional embeddings can be utilized for many downstream tasks such as simultaneous clustering, data visualization, and denoising. The proposed method is justified by rigorous theoretical analysis. Specifically, we show the consistency of our method in recovering the low-dimensional noiseless signals, and characterize the effects of the signal-to-noise ratios on the rates of convergence. Under a joint manifolds model framework, we establish the convergence of ultimate embeddings to the eigenfunctions of some newly introduced integral operators. These operators, referred to as duo-landmark integral operators, are defined by the convolutional kernel maps of some reproducing kernel Hilbert spaces (RKHSs). These RKHSs capture the either partially or entirely shared underlying low-dimensional nonlinear signal structures of the two datasets. Our numerical experiments and analyses of two single-cell omics datasets demonstrate the empirical advantages of the proposed method over existing methods in both embeddings and several downstream tasks.

stat.ML

Eigenvector distributions and optimal shrinkage estimators for large covariance and precision matrices

This paper focuses on investigating Stein's invariant shrinkage estimators for large sample covariance matrices and precision matrices in high-dimensional settings. We consider models that have nearly arbitrary population covariance matrices, including those with potential spikes. By imposing mild technical assumptions, we establish the asymptotic limits of the shrinkers for a wide range of loss functions. A key contribution of this work, enabling the derivation of the limits of the shrinkers, is a novel result concerning the asymptotic distributions of the non-spiked eigenvectors of the sample covariance matrices, which can be of independent interest.

math.ST

On the partial autocorrelation function for locally stationary time series: characterization, estimation and inference

For stationary time series, it is common to use the plots of partial autocorrelation function (PACF) or PACF-based tests to explore the temporal dependence structure of such processes. To our best knowledge, such analogs for non-stationary time series have not been fully established yet. In this paper, we fill this gap for locally stationary time series with short-range dependence. First, we characterize the PACF locally in the time domain and show that the $j$th PACF, denoted as $\rho_{j}(t),$ decays with $j$ whose rate is adaptive to the temporal dependence of the time series $\{x_{i,n}\}$. Second, at time $i,$ we justify that the PACF $\rho_j(i/n)$ can be efficiently approximated by the best linear prediction coefficients via the Yule-Walker's equations. This allows us to study the PACF via ordinary least squares (OLS) locally. Third, we show that the PACF is smooth in time for locally stationary time series. We use the sieve method with OLS to estimate $\rho_j(\cdot)$ and construct some statistics to test the PACFs and infer the structures of the time series. These tests generalize and modify those used for stationary time series. Finally, a multiplier bootstrap algorithm is proposed for practical implementation and an $\mathtt R$ package $\mathtt {Sie2nts}$ is provided to implement our algorithm. Numerical simulations and real data analysis also confirm usefulness of our results.

math.ST

Two sample test for covariance matrices in ultra-high dimension

In this paper, we propose a new test for testing the equality of two population covariance matrices in the ultra-high dimensional setting that the dimension is much larger than the sizes of both of the two samples. Our proposed methodology relies on a data splitting procedure and a comparison of a set of well selected eigenvalues of the sample covariance matrices on the split data sets. Compared to the existing methods, our methodology is adaptive in the sense that (i). it does not require specific assumption (e.g., comparable or balancing, etc.) on the sizes of two samples; (ii). it does not need quantitative or structural assumptions of the population covariance matrices; (iii). it does not need the parametric distributions or the detailed knowledge of the moments of the two populations. Theoretically, we establish the asymptotic distributions of the statistics used in our method and conduct the power analysis. We justify that our method is powerful under very weak alternatives. We conduct extensive numerical simulations and show that our method significantly outperforms the existing ones both in terms of size and power. Analysis of two real data sets is also carried out to demonstrate the usefulness and superior performance of our proposed methodology. An $\texttt{R}$ package $\texttt{UHDtst}$ is developed for easy implementation of our proposed methodology.

stat.ME

Global and local CLTs for linear spectral statistics of general sample covariance matrices when the dimension is much larger than the sample size with applications

In this paper, under the assumption that the dimension is much larger than the sample size, i.e., $p \asymp n^α, α>1,$ we consider the (unnormalized) sample covariance matrices $Q = Σ^{1/2} XX^*Σ^{1/2}$, where $X=(x_{ij})$ is a $p \times n$ random matrix with centered i.i.d entries whose variances are $(pn)^{-1/2}$, and $Σ$ is the deterministic population covariance matrix. We establish two classes of central limit theorems (CLTs) for the linear spectral statistics (LSS) for $Q,$ the global CLTs on the macroscopic scales and the local CLTs on the mesoscopic scales. We prove that the LSS converge to some Gaussian processes whose mean and covariance functions depending on $Σ$, the ratio $p/n$ and the test functions, can be identified explicitly on both macroscopic and mesoscopic scales. We also show that even though the global CLTs depend on the fourth cumulant of $x_{ij},$ the local CLTs do not. Based on these results, we propose two classes of statistics for testing the structures of $Σ,$ the global statistics and the local statistics, and analyze their superior power under general local alternatives. To our best knowledge, the local LSS testing statistics which do not rely on the fourth moment of $x_{ij},$ is used for the first time in hypothesis testing while the literature mostly uses the global statistics and requires the prior knowledge of the fourth cumulant. Numerical simulations also confirm the accuracy and powerfulness of our proposed statistics and illustrate better performance compared to the existing methods in the literature.

math.ST

Learning Low-Dimensional Nonlinear Structures from High-Dimensional Noisy Data: An Integral Operator Approach

We propose a kernel-spectral embedding algorithm for learning low-dimensional nonlinear structures from high-dimensional and noisy observations, where the datasets are assumed to be sampled from an intrinsically low-dimensional manifold and corrupted by high-dimensional noise. The algorithm employs an adaptive bandwidth selection procedure which does not rely on prior knowledge of the underlying manifold. The obtained low-dimensional embeddings can be further utilized for downstream purposes such as data visualization, clustering and prediction. Our method is theoretically justified and practically interpretable. Specifically, we establish the convergence of the final embeddings to their noiseless counterparts when the dimension and size of the samples are comparably large, and characterize the effect of the signal-to-noise ratio on the rate of convergence and phase transition. We also prove convergence of the embeddings to the eigenfunctions of an integral operator defined by the kernel map of some reproducing kernel Hilbert space capturing the underlying nonlinear structures. Numerical simulations and analysis of three real datasets show the superior empirical performance of the proposed method, compared to many existing methods, on learning various manifolds in diverse applications.

stat.ML

Auto-Regressive Approximations to Non-stationary Time Series, with Inference and Applications

Understanding the time-varying structure of complex temporal systems is one of the main challenges of modern time series analysis. In this paper, we show that every uniformly-positive-definite-in-covariance and sufficiently short-range dependent non-stationary and nonlinear time series can be well approximated globally by a white-noise-driven auto-regressive (AR) process of slowly diverging order. To our best knowledge, it is the first time such a structural approximation result is established for general classes of non-stationary time series. A high dimensional $\mathcal{L}^2$ test and an associated multiplier bootstrap procedure are proposed for the inference of the AR approximation coefficients. In particular, an adaptive stability test is proposed to check whether the AR approximation coefficients are time-varying, a frequently-encountered question for practitioners and researchers of time series. As an application, globally optimal short-term forecasting theory and methodology for a wide class of locally stationary time series are established via the method of sieves.

math.ST

Tracy-Widom distribution for the edge eigenvalues of elliptical model

In this paper, we study the largest eigenvalues of sample covariance matrices with elliptically distributed data. We consider the sample covariance matrix $Q=YY^*,$ where the data matrix $Y \in \mathbb{R}^{p \times n}$ contains i.i.d. $p$-dimensional observations $\mathbf{y}_i=ξ_iT\mathbf{u}_i,\;i=1,\dots,n.$ Here $\mathbf{u}_i$ is distributed on the unit sphere, $ξ_i \sim ξ$ is independent of $\mathbf{u}_i$ and $T^*T=Σ$ is some deterministic matrix. Under some mild regularity assumptions of $Σ,$ assuming $ξ^2$ has bounded support and certain proper behavior near its edge so that the limiting spectral distribution (LSD) of $Q$ has a square decay behavior near the spectral edge, we prove that the Tracy-Widom law holds for the largest eigenvalues of $Q$ when $p$ and $n$ are comparably large.

math.PR

Extreme eigenvalues of sample covariance matrices under generalized elliptical models with applications

We consider the extreme eigenvalues of the sample covariance matrix $Q=YY^*$ under the generalized elliptical model that $Y=Σ^{1/2}XD.$ Here $Σ$ is a bounded $p \times p$ positive definite deterministic matrix representing the population covariance structure, $X$ is a $p \times n$ random matrix containing either independent columns sampled from the unit sphere in $\mathbb{R}^p$ or i.i.d. centered entries with variance $n^{-1},$ and $D$ is a diagonal random matrix containing i.i.d. entries and independent of $X.$ Such a model finds important applications in statistics and machine learning. In this paper, assuming that $p$ and $n$ are comparably large, we prove that the extreme edge eigenvalues of $Q$ can have several types of distributions depending on $Σ$ and $D$ asymptotically. These distributions include: Gumbel, Fréchet, Weibull, Tracy-Widom, Gaussian and their mixtures. On the one hand, when the random variables in $D$ have unbounded support, the edge eigenvalues of $Q$ can have either Gumbel or Fréchet distribution depending on the tail decay property of $D.$ On the other hand, when the random variables in $D$ have bounded support, under some mild regularity assumptions on $Σ,$ the edge eigenvalues of $Q$ can exhibit Weibull, Tracy-Widom, Gaussian or their mixtures. Based on our theoretical results, we consider two important applications. First, we propose some statistics and procedure to detect and estimate the possible spikes for elliptically distributed data. Second, in the context of a factor model, by using the multiplier bootstrap procedure via selecting the weights in $D,$ we propose a new algorithm to infer and estimate the number of factors in the factor model. Numerical simulations also confirm the accuracy and powerfulness of our proposed methods and illustrate better performance compared to some existing methods in the literature.

stat.ME

Spiked multiplicative random matrices and principal components

In this paper, we study the eigenvalues and eigenvectors of the spiked invariant multiplicative models when the randomness is from Haar matrices. We establish the limits of the outlier eigenvalues $\widehatλ_i$ and the generalized components ($\langle \mathbf{v}, \widehat{\mathbf{u}}_i \rangle$ for any deterministic vector $\mathbb{v}$) of the outlier eigenvectors $\widehat{\mathbf{u}}_i$ with optimal convergence rates. Moreover, we prove that the non-outlier eigenvalues stick with those of the unspiked matrices and the non-outlier eigenvectors are delocalized. The results also hold near the so-called BBP transition and for degenerate spikes. On one hand, our results can be regarded as a refinement of the counterparts of [12] under additional regularity conditions. On the other hand, they can be viewed as an analog of [34] by replacing the random matrix with i.i.d. entries with Haar random matrix.

math.PR