SearcharxivSearch

arXiv subjects

Leheng Cai

Publications and source records attributed to Leheng Cai.

14 recordsLinked to original sources

Asymptotic Anytime-Valid Quantile Inference under Local Differential Privacy

Sequential quantile inference is difficult under local differential privacy because every record is randomized before reaching the analyst and the limiting quantile variance depends on an unknown density. We develop an online procedure that combines randomized response with dynamically chained parallel stochastic gradient descent (P-SGD). The resulting Polyak--Ruppert estimator admits a strong Gaussian approximation. A cross-chain quadratic statistic, computed entirely from private iterates, consistently estimates the limiting variance without a separate online density estimator. These results yield asymptotic confidence sequences and, under polynomial chain growth, asymptotic time-uniform coverage. Arm-wise constructions support locally private quantile best-arm identification, time-uniform simple-regret bounds, and sequential A/B tests of quantile treatment effects. Simulations and salary-data analyses illustrate the finite-sample behavior and practical use of the proposed methods.

stat.ME

Approximation Theorems for High-Dimensional Canonical U-Statistics: Gaussian Chaos and Phase Transition

We study simultaneous inference for maxima of canonical order-two $U$-statistics in high dimension. Degeneracy makes quadratic fluctuations leading, so ordinary Gaussian calibration can fail even after exact variance normalization. We show that the appropriate general target is a joint signed Gaussian quadratic chaos and establish a general approximation result that permits indefinite kernels. The general anti-concentration bound is too crude for high-dimensional inference, and we obtain sharper bounds under additional spectral structure. We also identify a phase transition from a non-Gaussian signed-chaos maximum to its covariance-matched Gaussian counterpart driven by the effective rank. For feasible inference, we propose a Gaussian multiplier bootstrap that avoid estimating eigensystems, and establish its validity. Two applications and extensive numerical simulations further illustrate the scope and practical performance of the proposed framework.

math.ST

Multicollinearity-agnostic feature screening for non-Euclidean responses: a factor adjusted approach

In high-dimensional settings, multicollinearity is a pervasive issue that can substantially impair the performance of feature screening methods based on marginal Fréchet regression. Feature screening for non-Euclidean responses becomes unreliable when ultrahigh-dimensional predictors suffer from multicollinearity, because feature-specific signals may be masked by shared latent factors. To mitigate this effect, we propose a Factor adjusted Fréchet sure independence screening procedure. The method first recovers latent common factors from the predictors and then evaluates each feature by the incremental Fréchet coefficient of determination contributed by its idiosyncratic component beyond the common factors. Under regularity conditions, we establish uniform approximation rates for the feasible screening utilities and prove the sure screening and sure ranking properties. Extensive numerical experiments provide compelling empirical support for the validity and effectiveness of our approach, particularly in scenarios with highly correlated covariates. We further illustrate the practical performance of our method through two representative non-Euclidean datasets: the ADNI dataset and the mortality dataset, both with distribution-valued responses.

stat.ME

Recursive Multiple Change Point Detection of Nonstationary Time Series: Instability Tests, Estimation and Confidence Intervals

We develop bootstrap-assisted robust binary segmentation (BARBS), a recursive binary segmentation method for multiple change point detection under general nonstationary temporal dynamics. A novel Gaussian multiplier bootstrap for the CUSUM statistics is proposed, offering robustness to complex dependence structures. Through meticulous calibration of the critical values at each stage of the recursion, BARBS ensures control of the Type I error under the null hypothesis of no change points. When change points are present, BARBS identifies the correct number of changes with a prespecified probability, and the resulting change point location estimators attain the same uniform consistency rate as classical binary segmentation. Building on this, we introduce second-stage refined estimators that achieve the optimal individual localization rate, and establish their asymptotic distributions and nearly optimal uniform localization rates under both fixed and vanishing jump magnitudes. Extensive numerical experiments across various settings confirm the robustness and superior performance of BARBS relative to existing approaches. To illustrate the practical relevance of the proposed methodology, we analyze U.S. inflation data, yielding change points that align with several documented macroeconomic episodes.

stat.ME

MATCH: Multiplier-Assisted Tests for Conditional Hypotheses in Non-Euclidean Data

We propose a new procedure MATCH (Multiplier-Assisted Tests for Conditional Hypotheses) to test whether the non-Euclidean data match the target model, which is a general framework for significance and specification testing in Fréchet regression. MATCH covers global significance, partial significance, and the adequacy of global Fréchet regression, providing a unified way to compare unrestricted conditional Fréchet means with restricted alternatives. One of the key challenges is that the ordinary held-out loss difference is first-order degenerate under the null: the oracle losses coincide, and plug-in statistics is dominated by nuisance estimation error. MATCH uses sample splitting and independent random multipliers on held-out losses to create a nondegenerate Gaussian leading term without residuals or tangent-space coordinates. To improve data use and stability, we further develop cross-fitted tests and repeated cross-fitting with p-value merging. We establish asymptotic null validity, consistency under fixed alternatives, and local power guarantees. Simulations for distributional, symmetric positive-definite (SPD) matrix-valued, and spherical responses support the theoretical findings, and applications to county-level household income distributions and North Atlantic tropical-cyclone locations demonstrate the practical use of the proposed tests.

stat.ME

Unified theory of testing relevant hypotheses in functional time series

In this paper, we develop a {\em unified} framework for testing relevant hypotheses in functional time series. The proposed approach accommodates one-sample, two-sample, and change point problems for contaminated observations under arbitrary sampling schemes. Combining B-spline estimation with self-normalization, we construct nuisance-parameter-free tests that bypass auxiliary estimation of long-run covariance functions and measurement-error variance functions. We establish asymptotic validity by exploiting a sequential Gaussian approximation for dependent random vectors of moderately high dimension, which leads to a pivotal limiting distribution. We also provide sufficient conditions for the non-degeneracy of the self-normalizer and establish consistent decision rules. A key theoretical finding is that the proposed tests detect \(n^{-1/2}\)-local alternatives under arbitrary sampling frequencies. This uncovers a sparse-to-dense phase transition distinct from those typically observed in functional data analysis: while the sampling frequency affects the asymptotic variance, the detection rate remains \(n^{-1/2}\), even in sparsely sampled regimes. We further study multiple change point alternatives and extend the theory to settings where consistent change point estimates are available. We also discuss the choice of self-normalizers, including the recently developed range-adjusted self-normalizer. Extensive simulations support the theoretical results, and applications to the AU.SHF implied volatility and traffic volume datasets demonstrate the practical utility of the proposed methods.

stat.ME

Asymptotic Anytime-Valid Inference for U-statistics

We study asymptotic anytime-valid confidence sequences for degree-two U-statistics under continuous monitoring. In the nondegenerate case, Hoeffding's projection reduces the problem to a time-uniform central limit theory for the partial sums of the first-order projection, while the canonical remainder is shown to be negligible under mild moment assumptions. A leave-one-out jackknife estimator then yields a fully data-driven procedure, leading to confidence sequences with asymptotic coverage guarantee for the parameter of interest. In the degenerate case, we show that the U-statistic is approximated by a centered quadratic Gaussian-chaos rather than by a simple Gaussian, which poses significant challenges for sequential inference. To address this issue, we novelly develop the Spectrally Allocated Gaussian-chaos Excursion (SAGE) boundary, and then provide plug-in implementations based on truncated spectrum estimation with consistency guarantees. The resulting widths can attain the expected time-uniform optimal rates: $\sqrt{\log\log n/n}$ in the nondegenerate regime and $\log\log n/n$ in the degenerate regime. Several widely used U-statistics are discussed within the proposed framework, and numerical experiments further support the validity of the derived theory.

math.ST

Strong Gaussian approximation for U-statistics in high dimensions and beyond

We establish a strong Gaussian approximation for high-dimensional non-degenerate U-statistics with diverging dimension. Under mild assumptions, we construct, on a sufficiently rich probability space, a Gaussian process that uniformly approximates the entire sequential U-statistic process. The approximation error is explicitly characterized and vanishes under polynomial growth of the dimension. The key technical contribution is a sharp martingale maximal inequality for completely degenerate U-statistics, combined with a high-dimensional strong approximation for independent sums. This coupling yields functional Gaussian limits without relying on $\mathcal{L}^\infty$-type bounds or bootstrap arguments. The theory is illustrated through three representative examples of U-statistics: the spatial Kendall's tau matrix, the multivariate Gini's mean difference, and the characteristic dispersion parameter. As applications, we derive Brownian bridge approximations for U-statistic-based change-point statistics and develop a self-normalized relevant testing procedure whose limiting distribution is fully pivotal. The framework naturally accommodates bounded kernels and therefore remains valid under heavy-tailed distributions. Overall, our results provide a unified probability-theoretic foundation for high-dimensional inference based on U-statistics.

math.ST

Federated Learning of Quantile Inference under Local Differential Privacy

In this paper, we investigate federated learning for quantile inference under local differential privacy (LDP). We propose an estimator based on local stochastic gradient descent (SGD), whose local gradients are perturbed via a randomized mechanism with global parameters, making the procedure tolerant of communication and storage constraints without compromising statistical efficiency. Although the quantile loss and its corresponding gradient do not satisfy standard smoothness conditions typically assumed in existing literature, we establish asymptotic normality for our estimator as well as a functional central limit theorem. The proposed method accommodates data heterogeneity and allows each server to operate with an individual privacy budget. Furthermore, we construct confidence intervals for the target value through a self-normalization approach, thereby circumventing the need to estimate additional nuisance parameters. Extensive numerical experiments and real data application validate the theoretical guarantees of the proposed methodology.

stat.ME

Adaptive adequacy testing of high-dimensional factor-augmented regression model

In this paper, we investigate the adequacy testing problem of high-dimensional factor-augmented regression model. Existing test procedures perform not well under dense alternatives. To address this critical issue, we introduce a novel quadratic-type test statistic which can efficiently detect dense alternative hypotheses. We further propose an adaptive test procedure to remain powerful under both sparse and dense alternative hypotheses. Theoretically, under the null hypothesis, we establish the asymptotic normality of the proposed quadratic-type test statistic and asymptotic independence of the newly introduced quadratic-type test statistic and a maximum-type test statistic. We also prove that our adaptive test procedure is powerful to detect signals under either sparse or dense alternative hypotheses. Simulation studies and an application to an FRED-MD macroeconomics dataset are carried out to illustrate the merits of our introduced procedures.

stat.ME

Statistical inference for high-dimensional convoluted rank regression

High-dimensional penalized rank regression is a powerful tool for modeling high-dimensional data due to its robustness and estimation efficiency. However, the non-smoothness of the rank loss brings great challenges to the computation. To solve this critical issue, high-dimensional convoluted rank regression has been recently proposed, introducing penalized convoluted rank regression estimators. However, these developed estimators cannot be directly used to make inference. In this paper, we investigate the statistical inference problem of high-dimensional convoluted rank regression. The use of U-statistic in convoluted rank loss function presents challenges for the analysis. We begin by establishing estimation error bounds of the penalized convoluted rank regression estimators under weaker conditions on the predictors. Building on this, we further introduce a debiased estimator and provide its Bahadur representation. Subsequently, a high-dimensional Gaussian approximation for the maximum deviation of the debiased estimator is derived, which allows us to construct simultaneous confidence intervals. For implementation, a novel bootstrap procedure is proposed and its theoretical validity is also established. Finally, simulation and real data analysis are conducted to illustrate the merits of our proposed methods.

stat.ME

From sparse to dense functional time series: phase transitions of detecting structural breaks and beyond

We develop a novel methodology for detecting abrupt break points in mean functions of functional time series, adaptable to arbitrary sampling schemes. By employing B-spline smoothing, we introduce $\mathcal L_{\infty}$ and $\mathcal L_2$ test statistics statistics based on a smoothed cumulative summation (CUMSUM) process, and derive the corresponding asymptotic distributions under the null and local alternative hypothesis, as well as the phase transition boundary from sparse to dense. We further establish the convergence rate of the proposed break point estimators and conduct statistical inference on the jump magnitude based on the estimated break point, also applicable across sparsely, semi-densely, and densely, observed random functions. Extensive numerical experiments validate the effectiveness of the proposed procedures. To illustrate the practical relevance, we apply the developed methods to analyze electricity price data and temperature data.

stat.ME

Test and Measure for Partial Mean Dependence Based on Machine Learning Methods

It is of importance to investigate the significance of a subset of covariates $W$ for the response $Y$ given covariates $Z$ in regression modeling. To this end, we propose a significance test for the partial mean independence problem based on machine learning methods and data splitting. The test statistic converges to the standard chi-squared distribution under the null hypothesis while it converges to a normal distribution under the fixed alternative hypothesis. Power enhancement and algorithm stability are also discussed. If the null hypothesis is rejected, we propose a partial Generalized Measure of Correlation (pGMC) to measure the partial mean dependence of $Y$ given $W$ after controlling for the nonlinear effect of $Z$. We present the appealing theoretical properties of the pGMC and establish the asymptotic normality of its estimator with the optimal root-$N$ convergence rate. Furthermore, the valid confidence interval for the pGMC is also derived. As an important special case when there are no conditional covariates $Z$, we introduce a new test of overall significance of covariates for the response in a model-free setting. Numerical studies and real data analysis are also conducted to compare with existing approaches and to demonstrate the validity and flexibility of our proposed procedures.

stat.ME

From Sparse to Dense Functional Data: Phase Transitions from a Simultaneous Inference Perspective

We aim to develop simultaneous inference tools for the mean function of functional data from sparse to dense. First, we derive a unified Gaussian approximation to construct simultaneous confidence bands of mean functions based on the B-spline estimator. Then, we investigate the conditions of phase transitions by decomposing the asymptotic variance of the approximated Gaussian process. As an extension, we also consider the orthogonal series estimator and show the corresponding conditions of phase transitions. Extensive simulation results strongly corroborate the theoretical results, and also illustrate the variation of the asymptotic distribution via the asymptotic variance decomposition we obtain. The developed method is further applied to body fat data and traffic data.

stat.ME