SearcharxivSearch

arXiv subjects

Wangli Xu

Publications and source records attributed to Wangli Xu.

13 recordsLinked to original sources

Random Projection Tests via Cauchy Combination for Two-Sample Mean

High-dimensional two-sample mean testing is challenging when the dimension exceeds the sample size. The random projection method proposed by Lopes et al. (2011) addresses this difficulty by mapping the data to a lower dimension space where Hotelling's $T^2$ statistic can be applied, while retaining useful covariance information and gaining power when the variables exhibit non-negligible covariance structure. However, single projection tests may be sensitive to the realized projection matrix, whereas existing multiple projection procedures often rely on resampling or simulation for calibration, with limited theoretical understanding. Moreover, projection-based Hotelling tests may be less sensitive to sparse mean differences. To address these limitations, we propose a Cauchy-combined random projection test (CRPT), which applies Hotelling's $T^2$ test after multiple independent random projections and combines the projected $p$-values through the Cauchy transformation. The proposed method retains the ability of random projection methods to incorporate covariance information while reducing reliance on any single projection. Under the Gaussian assumption, we establish the null tail behavior of the proposed statistic and further investigate its asymptotic power under suitable alternatives. To improve sensitivity to sparse alternatives, we further develop a power-enhanced version of CRPT. Simulation studies and real data analysis are conducted to examine the performance and practical applicability of the proposed procedures.

stat.ME

Deep Regression for Repeated Measurements under Covariate Shift

This paper studies nonparametric regression with repeated measurements when the response in the target domain is unobservable or costly to collect. We adopt a transfer learning framework that leverages a source domain with observable responses under covariate shift. The target regression function is estimated by correcting the distribution shift via the density ratio. We consider both known and unknown density ratio scenarios, which reflect different data available for nonparametric regression estimation. In both cases, we further address two settings: the uniformly bounded density ratio and the unbounded case with finite moment conditions. Under the unknown density ratio scenario, both the density ratio and the target regression function are estimated using rectified linear unit (ReLU) feedforward neural networks (FNNs), whereas under the known density ratio scenario, only the target regression function is estimated by ReLU FNNs. Theoretically, we establish non-asymptotic error bounds for the proposed estimators and prove that they achieve the minimax optimal convergence rate under the repeated measurements setting. Notably, we develop a novel approximation theory where the constants of the network parameters depend polynomially, rather than exponentially as in existing works, on the dimension, thereby mitigating the curse of dimensionality. Consequently, we derive sharper non-asymptotic bounds for the stochastic error. The finite sample performance of the proposed method is demonstrated through numerical simulations and a real data application.

stat.ME

SUP: An Inferable Private Multiple Testing Framework with Super Uniformity

Multiple testing is widely applied across scientific fields, particularly in genomic and health data analysis, where protecting sensitive personal information is imperative. However, developing private multiple testing algorithms for super uniform $p$-values remains an open question, as privacy mechanisms introduce intricate dependence among the peeled $p$-values and disrupt their super uniformity, complicating post-selection inference. To address this, we introduce a general Super Uniform Private (SUP) multiple testing framework with three key components. First, we develop a novel \( p \)-value transformation that is compatible with diverse privacy regimes while retaining the super uniformity. Next, a reversed peeling algorithm is designed to reduce privacy budgets while facilitating inference. Then, we provide diverse rejection thresholds that are privacy-parameter-free and tailored for different Type-I errors, including the family-wise error rate (FWER) and the false discovery rate (FDR). Building upon these, we advance adaptive techniques to determine the peeling number and boost thresholds. Theoretically, we propose a technique overcoming the post-selection obstacle to Type-I error control, quantify the privacy-induced power loss of SUP relative to its non-private counterpart, and demonstrate that SUP surpasses existing private methods in terms of power. The results of extensive simulations and a real data application validate our theories.

stat.ME

Varying-Coefficient Fr\'echet Regression

As a growing number of problems involve variables that are random objects, the development of models for such data has become increasingly important. This paper introduces a novel varying-coefficient Fr\'echet regression model that extends the classical varying-coefficient framework to accommodate random objects as responses. The proposed model provides a unified methodology for analyzing both Euclidean and non-Euclidean response variables. We develop a comprehensive estimation procedure that accommodates diverse predictor settings. Specifically, the model allows the effect-modifier variable U to be either Euclidean or non-Euclidean, while the predictors X are assumed to be Euclidean. Tailored estimation methods are provided for each scenario. To examine the asymptotic properties of the estimators, we introduce a smoothed version of the model and establish convergence rates through separate theoretical analyses of the bias and stochastic terms. The effectiveness and practical utility of the proposed methodology are demonstrated through extensive simulation studies and a real-data application.

stat.ME

An alternative measure for quantifying the heterogeneity in meta-analysis

Quantifying the heterogeneity is an important issue in meta-analysis, and among the existing measures, the $I^2$ statistic is most commonly used. In this paper, we first illustrate with a simple example that the $I^2$ statistic is heavily dependent on the study sample sizes, mainly because it is used to quantify the heterogeneity between the observed effect sizes. To reduce the influence of sample sizes, we introduce an alternative measure that aims to directly measure the heterogeneity between the study populations involved in the meta-analysis. We further propose a new estimator, namely the $I_A^2$ statistic, to estimate the newly defined measure of heterogeneity. For practical implementation, the exact formulas of the $I_A^2$ statistic are also derived under two common scenarios with the effect size as the mean difference (MD) or the standardized mean difference (SMD). Simulations and real data analysis demonstrate that the $I_A^2$ statistic provides an asymptotically unbiased estimator for the absolute heterogeneity between the study populations, and it is also independent of the study sample sizes as expected. To conclude, our newly defined $I_A^2$ statistic can be used as a supplemental measure of heterogeneity to monitor the situations where the study effect sizes are indeed similar with little biological difference. In such scenario, the fixed-effect model can be appropriate; nevertheless, when the sample sizes are sufficiently large, the $I^2$ statistic may still increase to 1 and subsequently suggest the random-effects model for meta-analysis.

stat.ME

Differentially Private Joint Independence Test

Identification of joint dependence among several random vectors plays an important role in many statistical applications, where the data may contain sensitive or confidential information. In this paper, we consider the $d$-variable Hilbert-Schmidt independence criterion (dHSIC) in the context of differential privacy. Given that the limiting distribution of the empirical estimate of dHSIC is a complicated Gaussian chaos, constructing tests in the non-private regime is typically based on permutation and bootstrap methods. To detect joint dependence under privacy constraints, we propose a dHSIC-based testing procedure employing a differentially private permutation methodology. We show that our method enjoys privacy guarantees, a valid level, and pointwise consistency, whereas the bootstrap counterpart suffers from inconsistent power. We further investigate the uniform power of the proposed test under the dHSIC and $L_2$ metrics, showing that the proposed test attains the minimax optimal power across different privacy regimes. As a byproduct, we show that the non-private permutation dHSIC test proposed in Pfister et al. (2018) is a special case of our differentially private permutation test, and our results also establish its pointwise and uniform power--thus resolving an open problem from that work. Both numerical simulations and real data analysis in causal inference suggest that our proposed test performs well empirically.

math.ST

Effective Positive Cauchy Combination Test

In the field of multiple hypothesis testing, combining p-values represents a fundamental statistical method. The Cauchy combination test (CCT) (Liu and Xie, 2020) excels among numerous methods for combining p-values with powerful and computationally efficient performance. However, large p-values may diminish the significance of testing, even extremely small p-values exist. We propose a novel approach named the positive Cauchy combination test (PCCT) to surmount this flaw. Building on the relationship between the PCCT and CCT methods, we obtain critical values by applying the Cauchy distribution to the PCCT statistic. We find, however, that the PCCT tends to be effective only when the significance level is substantially small or the test statistics are strongly correlated. Otherwise, it becomes challenging to control type I errors, a problem that also pertains to the CCT. Thanks to the theories of stable distributions and the generalized central limit theorem, we have demonstrated critical values under weak dependence, which effectively controls type I errors for any given significance level. For more general scenarios, we correct the test statistic using the generalized mean method, which can control the size under any dependence structure and cannot be further optimized. Our method exhibits excellent performance, as demonstrated through comprehensive simulation studies. We further validate the effectiveness of our proposed method by applying it to a genetic dataset.

stat.ME

Multifold Cross-Validation Model Averaging for Generalized Additive Partial Linear Models

Generalized additive partial linear models (GAPLMs) are appealing for model interpretation and prediction. However, for GAPLMs, the covariates and the degree of smoothing in the nonparametric parts are often difficult to determine in practice. To address this model selection uncertainty issue, we develop a computationally feasible model averaging (MA) procedure. The model weights are data-driven and selected based on multifold cross-validation (CV) (instead of leave-one-out) for computational saving. When all the candidate models are misspecified, we show that the proposed MA estimator for GAPLMs is asymptotically optimal in the sense of achieving the lowest possible Kullback-Leibler loss. In the other scenario where the candidate model set contains at least one correct model, the weights chosen by the multifold CV are asymptotically concentrated on the correct models. As a by-product, we propose a variable importance measure to quantify the importances of the predictors in GAPLMs based on the MA weights. It is shown to be able to asymptotically identify the variables in the true model. Moreover, when the number of candidate models is very large, a model screening method is provided. Numerical experiments show the superiority of the proposed MA method over some existing model averaging and selection methods.

stat.ME

Variable Importance Based Interaction Modeling with an Application on Initial Spread of COVID-19 in China

Interaction selection for linear regression models with both continuous and categorical predictors is useful in many fields of modern science, yet very challenging when the number of predictors is relatively large. Existing interaction selection methods focus on finding one optimal model. While attractive properties such as consistency and oracle property have been well established for such methods, they actually may perform poorly in terms of stability for high-dimensional data, and they do not typically deal with categorical predictors. In this paper, we introduce a variable importance based interaction modeling (VIBIM) procedure for learning interactions in a linear regression model with both continuous and categorical predictors. It delivers multiple strong candidate models with high stability and interpretability. Simulation studies demonstrate its good finite sample performance. We apply the VIBIM procedure to a Corona Virus Disease 2019 (COVID-19) data used in Tian et al. (2020) and measure the effects of relevant factors, including transmission control measures on the spread of COVID-19. We show that the VIBIM approach leads to better models in terms of interpretability, stability, reliability and prediction.

stat.ME

An approximate randomization test for high-dimensional two-sample Behrens-Fisher problem under arbitrary covariances

This paper is concerned with the problem of comparing the population means of two groups of independent observations. An approximate randomization test procedure based on the test statistic of Chen and Qin (2010) is proposed. The asymptotic behavior of the test statistic as well as the randomized statistic is studied under weak conditions. In our theoretical framework, observations are not assumed to be identically distributed even within groups. No condition on the eigenstructure of the covariance matrices is imposed. And the sample sizes of the two groups are allowed to be unbalanced. Under general conditions, all possible asymptotic distributions of the test statistic are obtained. We derive the asymptotic level and local power of the approximate randomization test procedure. Our theoretical results show that the proposed test procedure can adapt to all possible asymptotic distributions of the test statistic and always has correct test level asymptotically. Also, the proposed test procedure has good power behavior. Our numerical experiments show that the proposed test procedure has favorable performance compared with several alternative test procedures.

math.ST

Calibrated regression estimation using empirical likelihood under data fusion

Data analysis based on information from several sources is common in economic and biomedical studies. This setting is often referred to as the data fusion problem, which differs from traditional missing data problems since no complete data is observed for any subject. We consider a regression analysis when the outcome variable and some covariates are collected from two different sources. By leveraging the common variables observed in both data sets, doubly robust estimation procedures are proposed in the literature to protect against possible model misspecifications. However, they employ only a single propensity score model for the data fusion process and a single imputation model for the covariates available in one data set. It may be questionable to assume that either model is correctly specified in practice. We therefore propose an approach that calibrates multiple propensity score and imputation models to gain more protection based on empirical likelihood methods. The resulting estimator is consistent when any one of those models is correctly specified and is robust against extreme values of the fitted propensity scores. We also establish its asymptotic normality property and discuss the semiparametric estimation efficiency. Simulation studies show that the proposed estimator has substantial advantages over existing doubly robust estimators, and an assembled U.S. household expenditure data example is used for illustration.

stat.ME

Risk and return prediction for pricing portfolios of non-performing consumer credit

We design a system for risk-analyzing and pricing portfolios of non-performing consumer credit loans. The rapid development of credit lending business for consumers heightens the need for trading portfolios formed by overdue loans as a manner of risk transferring. However, the problem is nontrivial technically and related research is absent. We tackle the challenge by building a bottom-up architecture, in which we model the distribution of every single loan's repayment rate, followed by modeling the distribution of the portfolio's overall repayment rate. To address the technical issues encountered, we adopt the approaches of simultaneous quantile regression, R-copula, and Gaussian one-factor copula model. To our best knowledge, this is the first study that successfully adopts a bottom-up system for analyzing credit portfolio risks of consumer loans. We conduct experiments on a vast amount of data and prove that our methodology can be applied successfully in real business tasks.

q-fin.RM

Penalized Maximum Likelihood Estimator for Skew Normal Mixtures

Skew normal mixture models provide a more flexible framework than the popular normal mixtures for modelling heterogeneous data with asymmetric behaviors. Due to the unboundedness of likelihood function and the divergency of shape parameters, the maximum likelihood estimators of the parameters of interest are often not well defined, leading to dissatisfactory inferential process. We put forward a proposal to deal with these issues simultaneously in the context of penalizing the likelihood function. The resulting penalized maximum likelihood estimator is proved to be strongly consistent when the putative order of mixture is equal to or larger than the true one. We also provide penalized EM-type algorithms to compute penalized estimators. Finite sample performances are examined by simulations and real data applications and the comparison to the existing methods.

stat.ME