SearcharxivSearch

arXiv subjects

Wenzhi Yang

Publications and source records attributed to Wenzhi Yang.

6 recordsLinked to original sources

DROP: Distributionally Robust Optimization for Multi-task Learning in Graphical Models

Gaussian Graphical Models (GGMs) are widely used to infer conditional dependence structures in high-dimensional data. However, standard precision matrix estimators are highly sensitive to data contamination, such as extreme outliers and heavy-tailed noise. In this paper, we propose DROP (Distributionally Robust Optimization), a robust estimation method formulated within a multi-task nodewise regression framework. The proposed estimator enforces structural sparsity while resisting the influence of corrupted observations. Theoretically, we establish error bounds for the DROP estimator under general contamination. Through extensive high-dimensional simulations, we demonstrate that DROP consistently controls the rate of false positive edges and outperforms conventional non-robust estimators when data deviate from standard Gaussian assumptions. Furthermore, in a functional MRI (fMRI) application, DROP maintains a stable graph structure and preserves network modularity even when subjected to severe data perturbations, whereas competing methods yield excessively dense networks. To facilitate reproducible research, the DROP R package will be made publicly available on GitHub.

stat.AP

Consistent and powerful CUSUM change-point test for panel data with changes in variance

This paper investigates change-point of variance in panel data models with time series of $α$-mixing. Based on the cumulative sum (CUSUM) method and the individual differences, we construct a CUSUM test for panel data models to detect variance changes. Under the null hypothesis, we derive the limit distribution of this test, which can be used to detect the change-point of variance. Under the alternative hypothesis, the limit behavior of the CUSUM test is also derived. To validate the performance of the test, we conducted simulation analyses on with Gaussian and Gamma errors. The results demonstrate that this testing method significantly outperforms existing approaches, particularly in detecting sparse variance changes. Finally, we conducted a practical case study using panel data from the Shanghai Shenzhen CSI 300 Index Components. Not only did we successfully identify the change-points of variance, but we also delved deeper into the underlying economic drivers behind these changes.

stat.ME

Adaptive Kernel Regression for Constrained Route Alignment: Theory and Iterative Data Sharpening

Route alignment design in surveying and transportation engineering frequently involves fixed waypoint constraints, where a path must precisely traverse specific coordinates. While existing literature primarily relies on geometric optimization or control-theoretic spline frameworks, there is a lack of systematic statistical modeling approaches that balance global smoothness with exact point adherence. This paper proposes an Adaptive Nadaraya-Watson (ANW) kernel regression estimator designed to address the fixed waypoint problem. By incorporating waypoint-specific weight tuning parameters, the ANW estimator decouples global smoothing from local constraint satisfaction, avoiding the "jagged" artifacts common in naive local bandwidth-shrinking strategies. To further enhance estimation accuracy, we develop an iterative data sharpening algorithm that systematically reduces bias while maintaining the stability of the kernel framework. We establish the theoretical foundation for the ANW estimator by deriving its asymptotic bias and variance and proving its convergence properties under the internal constraint model. Numerical case studies in 1D and 2D trajectory planning demonstrate that the method effectively balances root mean square error (RMSE) and curvature smoothness. Finally, we validate the practical utility of the framework through empirical applications to railway and highway route planning. In sum, this work provides a stable, theoretically grounded, and computationally efficient solution for complex, constrained alignment design problems.

stat.ME

Q statistics in data depth: fundamental theory revisited and variants

Recently, data depth has been widely used to rank multivariate data. The study of the depth-based $Q$ statistic, originally proposed by Liu and Singh (1993), has become increasingly popular when it can be used as a quality index to differentiate between two samples. Based on the existing theoretical foundations, more and more variants have been developed for increasing power in the two sample test. However, the asymptotic expansion of the $Q$ statistic in the important foundation work of Zuo and He (2006) currently has an optimal rate $m^{-3/4}$ slower than the target $m^{-1}$, leading to limitations in higher-order expansions for developing more powerful tests. We revisit the existing assumptions and add two new plausible assumptions to obtain the target rate by applying a new proof method based on the Hoeffding decomposition and the Cox-Reid expansion. The aim of this paper is to rekindle interest in asymptotic data depth theory, to place Q-statistical inference on a firmer theoretical basis, to show its variants in current research, to open the door to the development of new theories for further variants requiring higher-order expansions, and to explore more of its potential applications.

math.ST

CLT for random quadratic forms based on sample means and sample covariance matrices

In this paper, we use the dimensional reduction technique to study the central limit theory (CLT) random quadratic forms based on sample means and sample covariance matrices. Specifically, we use a matrix denoted by $U_{p\times q}$, to map $q$-dimensional sample vectors to a $p$ dimensional subspace, where $q\geq p$ or $q\gg p$. Under the condition of $p/n\rightarrow 0$ as $(p,n)\rightarrow \infty$, we obtain the CLT of random quadratic forms for the sample means and sample covariance matrices.

math.ST

The Bahadur representation for sample quantiles under dependent sequence

On the one hand, we investigate the Bahadur representation for sample quantiles under $φ$-mixing sequence with $φ(n)=O(n^{-3})$ and obtain a rate as $O(n^{-\frac{3}{4}}\log n)$, $a.s.$. On the other hand, by relaxing the condition of mixing coefficients to $\sum\nolimits_{n=1}^\inftyφ^{1/2}(n)<\infty$, a rate $O(n^{-1/2}(\log n)^{1/2})$, $a.s.$, is also obtained.

math.ST