SearcharxivSearch

arXiv subjects

Long Feng

Publications and source records attributed to Long Feng.

At least 37 records · Page 2Linked to original sources

High-Dimensional Two-Sample Test for Elliptical Symmetry Distribution

We study the high-dimensional two-sample location problem under elliptical symmetry with arbitrary dependence in the scatter matrix. Existing spatial-sign procedures are attractive for heavy-tailed data, but their null calibration is tied to weakly dependent scatter matrices and their diagonal standardization does not, in general, recover the diagonal shape under strong dependence. We propose a new spatial-sign test based on coordinatewise pairwise-difference quantile scales. The new diagonal standardizer is location free, requires no positive moment condition on the radial variable, and estimates the diagonal of the elliptical shape up to a scalar specific to the sample, which disappears after spatial normalization. For the resulting full-sample statistic, we derive an explicit-rate stochastic expansion, establish a general weighted chi-square null distribution under arbitrary correlation structure, justify an empirical diagonal-deletion correction, and show that a Rademacher wild bootstrap consistently estimates the null law. The usual normal approximation appears only as a special case when no eigenvalue dominates.

stat.ME

High-Dimensional Tests for Elliptical Models via Radial--Directional Dependence

We develop high-dimensional goodness-of-fit tests for elliptical models by testing radial--directional independence after affine standardization. The method forms coordinatewise correlations between the log-radius and directional components, using a sum statistic for dense departures, a max statistic for sparse departures, and a Cauchy combination for adaptation. We derive oracle null limits, prove asymptotic independence of the sum and max components under both the null and a balanced local alternative, and establish validity of high-dimensional Hettmansperger--Randles plug-in standardization under explicit perturbation rates. Simulations and data analyses show stable size control, dense--sparse power complementarity, and interpretable coordinate-level diagnostics.

stat.ME

Sparse $K$-spatial-median clustering for high-dimensional data

We propose a robust clustering framework for high-dimensional data with heavy tails and a large fraction of irrelevant variables. The method replaces the mean updates of Lloyd's $K$-means with \emph{spatial medians} to enhance robustness. For the assignment step, it admits either a Euclidean rule for computational simplicity or a robust Mahalanobis-type metric constructed from the spatial sign covariance matrix to account for heterogeneous scales and feature dependence. To handle the $p \gg n$ regime, we further introduce a simple \emph{hard feature-exclusion} mechanism that removes weakly separating dimensions based on across-center dispersion, with the exclusion threshold selected automatically via a permutation-based Gap criterion. Simulation studies under correlated Gaussian and multivariate $t$ models demonstrate that the proposed approach provides competitive clustering accuracy and improved stability relative to $K$-means and sparse $K$-means baselines.

stat.ME

Testing Alpha in High-Dimensional Conditional Time-Varying Factor Models with Dependent Observations

This paper studies alpha testing in a high-dimensional conditional time-varying factor model with temporally dependent observations. Both factor loadings and alpha processes are allowed to vary smoothly over time, and the cross-sectional dimension may be comparable to or larger than the sample size. Using a B-spline sieve method, we develop a sum-type test for dense alternatives, a max-type test for sparse alternatives, and a Cauchy combination test for adaptive inference. On the theoretical side, we derive explicit stochastic expansions for the estimated average alphas, establish asymptotic normality of the sum statistic, and develop the extreme-value limit theory for the max statistic by showing its Gumbel convergence under temporal dependence together with the validity of block-bootstrap calibration. We further prove asymptotic independence between the sum and max statistics and thereby justify the Cauchy combination test. Simulation results demonstrate that the proposed procedures achieve satisfactory size control and competitive power across a wide range of dense and sparse alternatives. An empirical application further illustrates the usefulness of the proposed methods in testing asset-pricing models with time-varying structure.

stat.ME

CORA: Conformal Risk-Controlled Agents for Safeguarded Mobile GUI Automation

Graphical user interface (GUI) agents powered by vision language models (VLMs) are rapidly moving from passive assistance to autonomous operation. However, this unrestricted action space exposes users to severe and irreversible financial, privacy or social harm. Existing safeguards rely on prompt engineering, brittle heuristics and VLM-as-critic lack formal verification and user-tunable guarantees. We propose CORA (COnformal Risk-controlled GUI Agent), a post-policy, pre-action safeguarding framework that provides statistical guarantees on harmful executed actions. CORA reformulates safety as selective action execution: we train a Guardian model to estimate action-conditional risk for each proposed step. Rather than thresholding raw scores, we leverage Conformal Risk Control to calibrate an execute/abstain boundary that satisfies a user-specified risk budget and route rejected actions to a trainable Diagnostician model, which performs multimodal reasoning over rejected actions to recommend interventions (e.g., confirm, reflect, or abort) to minimize user burden. A Goal-Lock mechanism anchors assessment to a clarified, frozen user intent to resist visual injection attacks. To rigorously evaluate this paradigm, we introduce Phone-Harm, a new benchmark of mobile safety violations with step-level harm labels under real-world settings. Experiments on Phone-Harm and public benchmarks against diverse baselines validate that CORA improves the safety--helpfulness--interruption Pareto frontier, offering a practical, statistically grounded safety paradigm for autonomous GUI execution. Code and benchmark are available at cora-agent.github.io.

cs.LG

Maximum-of-Differences Test for Comparing Multivariate K-Sample Distributions

Comparing $K$-sample distributions is a fundamental problem in data science that arises in a wide variety of fields and applications. In this article, we introduce a maximum-of-differences approach to make such comparisons. Specifically, we first calculate the pairwise distances from the pooled observations of the $K$ samples. We then define the two observations as connected if their distance is less than a pre-specified threshold value. For each observation, we next calculate the ``within" and the ``between" probabilities associated with these two types of connections for the given observation, i.e., with other observations within the same sample and between the given observation and the observations in other samples. Subsequently, we propose a maximum-of-differences (MOD) test that finds the maximum value among the standardized squared differences between the ``within" and the ``between" probabilities of all observations. Accordingly, the proposed test is not only applicable to multivariate data with $K$ samples, but can also be extended to multivariate regression models. Furthermore, we obtain the covariance-adjusted (CA) version of the MOD (CA-MOD) test, which converges to the Type I extreme value distribution under some conditions. Moreover, we demonstrate the asymptotic properties of the two tests under both the null and alternative hypotheses. The performance and usefulness of the tests are illustrated via simulation studies and real examples.

stat.ME

High Dimensional Bootstrap and Asymptotic Expansion for the $k$-th Largest Coordinate

We study bootstrap inference for the $k$th largest coordinate of a normalized sum of independent high-dimensional random vectors. Existing second-order theory for maxima does not directly extend to order statistics, because the event $\{T_{n,[k]}\le t\}$ is not a rectangle and its local structure is governed by exceedance counts rather than by a single boundary. We develop an approach based on factorial moments and weighted inclusion--exclusion that reduces the problem to a collection of rare-orthant probabilities and allows high-dimensional Edgeworth and Cornish--Fisher expansions to be transferred to the order-statistic setting. Under moment, variance, and weak-dependence conditions, we derive a second-order coverage expansion for wild-bootstrap critical values of the $k$th order statistic. In particular, a third-moment matching wild bootstrap achieves coverage error of order $n^{-1}$ up to logarithmic factors, and the same second-order accuracy is obtained for a prepivoted double wild bootstrap. We also show that the maximal-correlation condition can be replaced by a stationary Gaussian exponential-mixing assumption at the price of an explicit dependence remainder $r_d$, and this remainder can itself be of order $n^{-1}$ when the dimension is sufficiently large relative to the sample size. These results extend recent second-order Gaussian and bootstrap approximation theory from maxima to the $k$th order statistic in high dimension.

math.ST

Rank-Based Sparse Regression in Principal Components Space under Measurement Error

We study high-dimensional regression in principal components space when the predictors are observed with additive measurement error and the response errors may be heavy-tailed. The starting point is the $\ell_1$-penalized principal-components estimator of Song and Zou (2026), which enjoys a blessing-of-dimensionality phenomenon under predictor contamination but senstive for heavy-tailed data or outliers. We replace the squared loss by a Wilcoxon-type rank loss and then apply a one-step adaptive reweighting scheme to reduce the shrinkage bias of the initial $\ell_1$ fit. The resulting procedure combines robustness to heavy-tailed response errors with the contamination geometry induced by the empirical principal-components basis. Our main theorem gives a prediction bound for the fixed-$λ$ second-stage fitted mean. Simulations show that the rank-based procedure is competitive under Gaussian noise and substantially more stable under heavy-tailed errors, especially when predictor contamination is present.

stat.ME

High dimensional alpha test for linear factor pricing model with $L_q$-norm

We consider testing zero pricing errors in high-dimensional linear factor pricing models. Existing methods are mainly based on either an $L_2$ statistic, which is effective under dense alternatives, or an $L_\infty$ statistic, which is powerful under very sparse alternatives. To bridge these two regimes, we develop a class of $L_q$-based tests for finite $q$, including the practically useful $L_4$ and $L_6$ cases. We show that larger $q$ leads to greater sensitivity to sparse alternatives. We further establish the asymptotic independence between the $L_\infty$ statistic and the $L_q$ statistic for any finite $q$, which motivates a Cauchy combination test that adapts to a broad range of sparsity levels. Simulation studies and a real-data analysis show that the proposed methods are more robust to the unknown sparsity of the alternative and can outperform existing procedures in finite samples.

stat.ME

Difference-Based High-Dimensional Long-Run Covariance Matrix Estimation for Mean-shift Time Series

We consider estimation of high-dimensional long-run covariance matrices for time series with nonconstant means, a setting in which conventional estimators can be severely biased. To address this difficulty, we propose a difference-based initial estimator that is robust to a broad class of mean variations, and combine it with hard thresholding, soft thresholding, and tapering to obtain sparse long-run covariance estimators for high-dimensional data. We derive convergence rates for the resulting estimators under general temporal dependence and time-varying mean structures, showing explicitly how the rates depend on covariance sparsity, mean variation, dimension, and sample size. Numerical experiments show that the proposed methods perform favorably in high dimensions, especially when the mean evolves over time.

stat.ME

Conditional Rank-Rank Regression via Deep Conditional Transformation Models

Intergenerational mobility quantifies the transmission of socio-economic outcomes from parents to children. While rank-rank regression (RRR) is standard, adding covariates directly (RRRX) often yields parameters with unclear interpretation. Conditional rank-rank regression (CRRR) resolves this by using covariate-adjusted (conditional) ranks to measure within-group mobility. We improve and extend CRRR by estimating conditional ranks with a deep conditional transformation model (DCTM) and cross-fitting, enabling end-to-end conditional distribution learning with structural constraints and strong performance under nonlinearity, high-order interactions, and discrete ordered outcomes where the distributional regression used in traditional CRRR may be cumbersome or prone to misconfiguration. We further extend CRRR to discrete outcomes via an $ω$-indexed conditional-rank definition and study sensitivity to $ω$. For continuous outcomes, we establish an asymptotic theory for the proposed estimators and verify the validity of exchangeable bootstrap inference. Simulations across simple/complex continuous and discrete ordered designs show clear accuracy gains in challenging settings. Finally, we apply our method to two empirical studies, revealing substantial within-group persistence in U.S. income and pronounced gender differences in educational mobility in India.

stat.ME

Semi-Supervised Generative Learning via Latent Space Distribution Matching

We introduce Latent Space Distribution Matching (LSDM), a novel framework for semi-supervised generative modeling of conditional distributions. LSDM operates in two stages: (i) learning a low-dimensional latent space from both paired and unpaired data, and (ii) performing joint distribution matching in this space via the 1-Wasserstein distance, using only paired data. This two-step approach minimizes an upper bound on the 1-Wasserstein distance between joint distributions, reducing reliance on scarce paired samples while enabling fast one-step generation. Theoretically, we establish non-asymptotic error bounds and demonstrate a key benefit of unpaired data: enhanced geometric fidelity in generated outputs. Furthermore, by extending the scope of its two core steps, LSDM provides a coherent statistical perspective that connects to a broad class of latent-space approaches. Notably, Latent Diffusion Models (LDMs) can be viewed as a variant of LSDM, in which joint distribution matching is achieved indirectly via score matching. Consequently, our results also provide theoretical insights into the consistency of LDMs. Empirical evaluations on real-world image tasks, including class-conditional generation and image super-resolution, demonstrate the effectiveness of LSDM in leveraging unpaired data to enhance generation quality.

stat.ML

Note on High Dimensional Spatial-Sign Test for One Sample Problem

We revisit the null distribution of the high-dimensional spatial-sign test of Wang et al. (2015) under mild structural assumptions on the scatter matrix. We show that the standardized test statistic converges to a non-Gaussian limit, characterized as a mixture of a normal component and a weighted chi-square component. To facilitate practical implementation, we propose a wild bootstrap procedure for computing critical values and establish its asymptotic validity. Numerical experiments demonstrate that the proposed bootstrap test delivers accurate size control across a wide range of dependence settings and dimension-sample-size regimes.

stat.ME

Deep Kronecker Network

We propose Deep Kronecker Network (DKN), a novel framework designed for analyzing medical imaging data, such as MRI, fMRI, CT, etc. Medical imaging data is different from general images in at least two aspects: i) sample size is usually much more limited, ii) model interpretation is more of a concern compared to outcome prediction. Due to its unique nature, general methods, such as convolutional neural network (CNN), are difficult to be directly applied. As such, we propose DKN, that is able to i) adapt to low sample size limitation, ii) provide desired model interpretation, and iii) achieve the prediction power as CNN. The DKN is general in the sense that it not only works for both matrix and (high-order) tensor represented image data, but also could be applied to both discrete and continuous outcomes. The DKN is built on a Kronecker product structure and implicitly imposes a piecewise smooth property on coefficients. Moreover, the Kronecker structure can be written into a convolutional form, so DKN also resembles a CNN, particularly, a fully convolutional network (FCN). Furthermore, we prove that with an alternating minimization algorithm, the solutions of DKN are guaranteed to converge to the truth geometrically even if the objective function is highly nonconvex. Interestingly, the DKN is also highly connected to the tensor regression framework proposed by Zhou et al. (2010), where a CANDECOMP/PARAFAC (CP) low-rank structure is imposed on tensor coefficients. Finally, we conduct both classification and regression analyses using real MRI data from the Alzheimer's Disease Neuroimaging Initiative (ADNI) to demonstrate the effectiveness of DKN.

stat.ML

High dimensional matrix estimation through elliptical factor models

Elliptical factor models play a central role in modern high-dimensional data analysis, particularly due to their ability to capture heavy-tailed and heterogeneous dependence structures. Within this framework, Tyler's M-estimator (Tyler, 1987a) enjoys several optimality properties and robustness advantages. In this paper, we develop high-dimensional scatter matrix, covariance matrix and precision matrix estimators grounded in Tyler's M-estimation. We first adapt the Principal Orthogonal complEment Thresholding (POET) framework (Fan et al., 2013) by incorporating the spatial-sign covariance matrix as an effective initial estimator. Building on this idea, we further propose a direct extension of POET tailored for Tyler's M-estimation, referred to as the POET-TME method. We establish the consistency rates for the resulting estimators under elliptical factor models. Comprehensive simulation studies and a real data application illustrate the superior performance of POET-TME, especially in the presence of heavy-tailed distributions, demonstrating the practical value of our methodological contributions.

stat.ME

A Nonparametric Statistics Approach to Feature Selection in Deep Neural Networks with Theoretical Guarantees

This paper tackles the problem of feature selection in a highly challenging setting: $\mathbb{E}(y | \boldsymbol{x}) = G(\boldsymbol{x}_{\mathcal{S}_0})$, where $\mathcal{S}_0$ is the set of relevant features and $G$ is an unknown, potentially nonlinear function subject to mild smoothness conditions. Our approach begins with feature selection in deep neural networks, then generalizes the results to H{ö}lder smooth functions by exploiting the strong approximation capabilities of neural networks. Unlike conventional optimization-based deep learning methods, we reformulate neural networks as index models and estimate $\mathcal{S}_0$ using the second-order Stein's formula. This gradient-descent-free strategy guarantees feature selection consistency with a sample size requirement of $n = Ω(p^2)$, where $p$ is the feature dimension. To handle high-dimensional scenarios, we further introduce a screening-and-selection mechanism that achieves nonlinear selection consistency when $n = Ω(s \log p)$, with $s$ representing the sparsity level. Additionally, we refit a neural network on the selected features for prediction and establish performance guarantees under a relaxed sparsity assumption. Extensive simulations and real-data analyses demonstrate the strong performance of our method even in the presence of complex feature interactions.

stat.ML

High dimensional Mean Test for Temporal Dependent Data

This paper proposes a novel test method for high-dimensional mean testing regard for the temporal dependent data. Comparison to existing methods, we establish the asymptotic normality of the test statistic without relying on restrictive assumptions, such as Gaussian distribution or M-dependence. Importantly, our theoretical framework holds potential for extension to other high-dimensional problems involving temporal dependent data. Additionally, our method offers significantly reduced computational complexity, making it more practical for large-scale applications. Simulation studies further demonstrate the computational advantages and performance improvements of our test.

stat.ME

Enterprise Profit Prediction Using Multiple Data Sources with Missing Values through Vertical Federated Learning

Small and medium-sized enterprises (SMEs) play a crucial role in driving economic growth. Monitoring their financial performance and discovering relevant covariates are essential for risk assessment, business planning, and policy formulation. This paper focuses on predicting profits for SMEs. Two major challenges are faced in this study: 1) SMEs data are stored across different institutions, and centralized analysis is restricted due to data security concerns; 2) data from various institutions contain different levels of missing values, resulting in a complex missingness issue. To tackle these issues, we introduce an innovative approach named Vertical Federated Expectation Maximization (VFEM), designed for federated learning under a missing data scenario. We embed a new EM algorithm into VFEM to address complex missing patterns when full dataset access is unfeasible. Furthermore, we establish the linear convergence rate for the VFEM and establish a statistical inference framework, enabling covariates to influence assessment and enhancing model interpretability. Extensive simulation studies are conducted to validate its finite sample performance. Finally, we thoroughly investigate a real-life profit prediction problem for SMEs using VFEM. Our findings demonstrate that VFEM provides a promising solution for addressing data isolation and missing values, ultimately improving the understanding of SMEs' financial performance.

stat.ME