SearcharxivSearch

arXiv subjects

Qizhai Li

Publications and source records attributed to Qizhai Li.

12 recordsLinked to original sources

Inverse regression for causal inference with multiple outcomes

With multiple outcomes in empirical research, a common strategy is to define a composite outcome as a weighted average of the original outcomes. However, the choices of weights are often subjective and can be controversial. We propose an inverse regression strategy for causal inference with multiple outcomes. The key idea is to regress the treatment on the outcomes, which is the inverse of the standard regression of the outcomes on the treatment. Although this strategy is simple and even counterintuitive, it has several advantages. First, testing for zero coefficients of the outcomes is equivalent to testing for zero treatment effects, even though the inverse regression is deemed misspecified. Second, the coefficients of the outcomes provide data-driven weights for defining a composite outcome. Interestingly, these weights maximize the standardized effect size, a measure commonly used for univariate outcomes in empirical studies. We also discuss the associated inference issues. Third, this strategy is applicable to general study designs. We illustrate the theory in both randomized experiments and observational studies. We implement the proposed method in our R package invreg, available at https://github.com/zhangwei0125/invreg.

stat.ME

Gaussian Multiplier Bootstrap Procedure for the $k$th Largest Coordinate of High-Dimensional Statistics

We consider the problem of Gaussian multiplier bootstrap procedures for the $k$th largest statistics and functions of the top $k$ order statistics, which are commonly encountered in high-dimensional statistical inference. Such a problem has been studied previously for $k=1$ (i.e., maxima). However, in many applications, a general $k$ ($k\geq 1$) is of great interest. We provide the upper bounds for the errors between Gaussian approximations and Gaussian multiplier approximations. The dimension $p$ is allowed to be larger than the sample size $n$. The effectiveness of the proposed methods is demonstrated via the computer numerical results and a real-world data analysis.

math.ST

Limit theorems of Azadkia-Chatterjee's conditional graph correlation

Inferring the strength of conditional dependence and testing conditional independence are fundamental problems in statistics. A recent breakthrough by Azadkia and Chatterjee introduced, for the first time, a conditional dependence measure that equals $0$ if and only if the variables under study are conditionally independent, and equals $1$ if and only if they are conditionally perfectly dependent. They further proposed a computationally efficient and strongly consistent estimator, $T_n$, based on an ingenious use of ranks and nearest neighbors. Despite these attractive features, the asymptotic theory of $T_n$ has remained largely undeveloped. This paper closes that gap. We prove that, under general dependence, $T_n$ is asymptotically normal and its limiting variance admits a closed form. We also construct consistent variance estimators that are computationally efficient and implementable in $O(n\log n)$ time. Taken together with existing bias-correction methods, these results provide a complete inferential theory for $T_n$.

math.ST

Principled Inference in Dense High-Dimensional Linear Models via Local Conditional Sparsity

High-dimensional inference methods often rely on coefficient sparsity, an assumption that can be restrictive when signals are dense but individually weak. In such settings, valid inference may still be possible if the covariates exhibit sparse conditional dependence. Motivated by this observation, we propose Neighborhood-Localized Nested Regression (NLNR), a framework for coordinatewise inference in high-dimensional linear models with potentially dense coefficients. The central idea is to localize inference for a target coefficient to a low-dimensional working regression determined by a Sparse Conditional Neighborhood (SCN) of the target covariate. Specifically, for a given covariate, we estimate its SCN through nodewise $\ell_1$-penalized regression and then fit a regression using only the target covariate and its estimated neighborhood. Under suitable regularity conditions, we establish consistency and asymptotic normality of the resulting estimator. Building on this inferential reduction principle, we further develop a thresholding-based screening procedure with theoretical guarantees and a boosting variant that augments the working model with additional response-relevant covariates to improve finite-sample performance. Extensive simulations and an application to the CCLE dataset demonstrate favorable empirical performance.

stat.ME

Gaussian Approximations for the $k$th coordinate of sums of random vectors

We consider the problem of Gaussian approximation for the $κ$th coordinate of a sum of high-dimensional random vectors. Such a problem has been studied previously for $κ=1$ (i.e., maxima). However, in many applications, a general $κ\geq1$ is of great interest, which is addressed in this paper. We make four contributions: 1) we first show that the distribution of the $κ$th coordinate of a sum of random vectors, $\boldsymbol{X}= (X_{1},\cdots,X_{p})^{\sf T}= n^{-1/2}\sum_{i=1}^n \boldsymbol{x}_{i}$, can be approximated by that of Gaussian random vectors and derive their Kolmogorov's distributional difference bound; 2) we provide the theoretical justification for estimating the distribution of the $κ$th coordinate of a sum of random vectors using a Gaussian multiplier procedure, which multiplies the original vectors with i.i.d. standard Gaussian random variables; 3) we extend the Gaussian approximation result and Gaussian multiplier bootstrap procedure to a more general case where $κ$ diverges; 4) we further consider the Gaussian approximation for a square sum of the first $d$ largest coordinates of $\boldsymbol{X}$. All these results allow the dimension $p$ of random vectors to be as large as or much larger than the sample size $n$.

math.ST

Tight Generalization Error Bounds for Stochastic Gradient Descent in Non-convex Learning

Stochastic Gradient Descent (SGD) is fundamental for training deep neural networks, especially in non-convex settings. Understanding SGD's generalization properties is crucial for ensuring robust model performance on unseen data. In this paper, we analyze the generalization error bounds of SGD for non-convex learning by introducing the Type II perturbed SGD (T2pm-SGD), which accommodates both sub-Gaussian and bounded loss functions. The generalization error bound is decomposed into two components: the trajectory term and the flatness term. Our analysis improves the trajectory term to $O(n^{-1})$, significantly enhancing the previous $O((nb)^{-1/2})$ bound for bounded losses, where n is the number of training samples and b is the batch size. By selecting an optimal variance for the perturbation noise, the overall bound is further refined to $O(n^{-2/3})$. For sub-Gaussian loss functions, a tighter trajectory term is also achieved. In both cases, the flatness term remains stable across iterations and is smaller than those reported in previous literature, which increase with iterations. This stability, ensured by T2pm-SGD, leads to tighter generalization error bounds for both loss function types. Our theoretical results are validated through extensive experiments on benchmark datasets, including MNIST and CIFAR-10, demonstrating the effectiveness of T2pm-SGD in establishing tighter generalization bounds.

stat.ML

Functional knockoffs selection with applications to functional data analysis in high dimensions

The knockoffs is a recently proposed powerful framework that effectively controls the false discovery rate (FDR) for variable selection. However, none of the existing knockoff solutions are directly suited to handle multivariate or high-dimensional functional data, which has become increasingly prevalent in various scientific applications. In this paper, we propose a novel functional model-X knockoffs selection framework tailored to sparse high-dimensional functional models, and show that our proposal can achieve the effective FDR control for any sample size. Furthermore, we illustrate the proposed functional model-X knockoffs selection procedure along with the associated theoretical guarantees for both FDR control and asymptotic power using examples of commonly adopted functional linear additive regression models and the functional graphical model. In the construction of functional knockoffs, we integrate essential components including the correlation operator matrix, the Karhunen-Loève expansion, and semidefinite programming, and develop executable algorithms. We demonstrate the superiority of our proposed methods over the competitors through both extensive simulations and the analysis of two brain imaging datasets.

stat.ME

A family of Chatterjee's correlation coefficients and their properties

Quantifying the strength of functional dependence between random scalars $X$ and $Y$ is an important statistical problem. While many existing correlation coefficients excel in identifying linear or monotone functional dependence, they fall short in capturing general non-monotone functional relationships. In response, we propose a family of correlation coefficients $ξ^{(h,F)}_n$, characterized by a continuous bivariate function $h$ and a cdf function $F$. By offering a range of selections for $h$ and $F$, $ξ^{(h,F)}_n$ encompasses a diverse class of novel correlation coefficients, while also incorporates the Chatterjee's correlation coefficient (Chatterjee, 2021) as a special case. We prove that $ξ^{(h,F)}_n$ converges almost surely to a deterministic limit $ξ^{(h,F)}$ as sample size $n$ approaches infinity. In addition, under appropriate conditions imposed on $h$ and $F$, the limit $ξ^{(h,F)}$ satisfies the three appealing properties: (P1). it belongs to the range of $[0,1]$; (P2). it equals 1 if and only if $Y$ is a measurable function of $X$; and (P3). it equals 0 if and only if $Y$ is independent of $X$. As amplified by our numerical experiments, our proposals provide practitioners with a variety of options to choose the most suitable correlation coefficient tailored to their specific practical needs.

stat.ME

The Cauchy Combination Test under Arbitrary Dependence Structures

Aggregating multiple effects is often encountered in large-scale data analysis where the fraction of significant effects is generally small. Many existing methods cannot handle it effectively because of lack of computational accuracy for small p-values. The Cauchy combination test (abbreviated as CCT) ( J Am Statist Assoc, 2020, 115(529):393-402) is a powerful and computational effective test to aggregate individual $p$-values under arbitrary correlation structures. This work revisits CCT and shows three key contributions including that (i) the tail probability of CCT can be well approximated by a standard Cauchy distribution under much more relaxed conditions placed on individual p-values instead of the original test statistics; (ii) the relaxation conditions are shown to be satisfied for many popular copulas formulating bivariate distributions; (iii) the power of CCT is no less than that of the minimum-type test as the number of tests goes to infinity with some regular conditions. These results further broaden the theories and applications of CCT. The simulation results verify the theoretic results and the performance of CCT is further evaluated with data from a prostate cancer study.

stat.ME

Distance-based regression analysis for measuring associations

Distance-based regression model, as a nonparametric multivariate method, has been widely used to detect the association between variations in a distance or dissimilarity matrix for outcomes and predictor variables of interest in genetic association studies, genomic analyses, and many other research areas. Based on it, a pseudo-$F$ statistic which partitions the variation in distance matrices is often constructed to achieve the aim. To the best of our knowledge, the statistical properties of the pseudo-$F$ statistic has not yet been well established in the literature. To fill this gap, we study the asymptotic null distribution of the pseudo-$F$ statistic and show that it is asymptotically equivalent to a mixture of chi-squared random variables. Given that the pseudo-$F$ test statistic has unsatisfactory power when the correlations of the response variables are large, we propose a square-root $F$-type test statistic which replaces the similarity matrix with its square root. The asymptotic null distribution of the new test statistic and power of both tests are also investigated. Simulation studies are conducted to validate the asymptotic distributions of the tests and demonstrate that the proposed test has more robust power than the pseudo-$F$ test. Both test statistics are exemplified with a gene expression dataset for a prostate cancer pathway. Keywords: Asymptotic distribution, Chi-squared-type mixture, Nonparametric test, Pseudo-$F$ test, Similarity matrix.

math.ST

An Imputation-Consistency Algorithm for High-Dimensional Missing Data Problems and Beyond

Missing data are frequently encountered in high-dimensional problems, but they are usually difficult to deal with using standard algorithms, such as the expectation-maximization (EM) algorithm and its variants. To tackle this difficulty, some problem-specific algorithms have been developed in the literature, but there still lacks a general algorithm. This work is to fill the gap: we propose a general algorithm for high-dimensional missing data problems. The proposed algorithm works by iterating between an imputation step and a consistency step. At the imputation step, the missing data are imputed conditional on the observed data and the current estimate of parameters; and at the consistency step, a consistent estimate is found for the minimizer of a Kullback-Leibler divergence defined on the pseudo-complete data. For high dimensional problems, the consistent estimate can be found under sparsity constraints. The consistency of the averaged estimate for the true parameter can be established under quite general conditions. The proposed algorithm is illustrated using high-dimensional Gaussian graphical models, high-dimensional variable selection, and a random coefficient model.

stat.ME