SearcharxivSearch

arXiv subjects

Shihua Luo

Publications and source records attributed to Shihua Luo.

5 recordsLinked to original sources

Deep Neural Networks for Doubly Robust Estimation with Nonprobability Survey Samples

Integrating probability and nonprobability survey samples is an important problem in modern survey sampling. Nonprobability samples often contain rich outcome information but may lack population representativeness, whereas probability samples provide design-based auxiliary information but may not contain the study variable. We propose a deep neural network (DNN)-assisted doubly robust framework for estimating the finite population mean from these two data sources. The proposed method models the logit sampling score for the nonprobability sample as an unknown nonparametric function and estimates it by maximizing a pseudo-likelihood that combines information from the nonprobability sample and a reference probability sample. The DNN parameters are optimized using the ADAM algorithm. The resulting DNN-estimated sampling scores are incorporated into a DNN-assisted inverse-probability weighted estimator and a deep doubly robust estimator. We establish consistency and convergence rates under regularity conditions and evaluate the finite-sample performance of the proposed estimators through simulation studies and an empirical application using Pew Research Center and Behavioral Risk Factor Surveillance System data. The results suggest that the proposed estimators can improve robustness to parametric propensity-score misspecification, especially when the true selection mechanism is nonlinear.

math.ST

Energy-Based Model for Accurate Estimation of Shapley Values in Feature Attribution

Shapley value is a widely used tool in explainable artificial intelligence (XAI), as it provides a principled way to attribute contributions of input features to model outputs. However, estimation of Shapley value requires capturing conditional dependencies among all feature combinations, which poses significant challenges in complex data environments. In this article, EmSHAP (Energy-based model for Shapley value estimation), an accurate Shapley value estimation method, is proposed to estimate the expectation of Shapley contribution function under the arbitrary subset of features given the rest. By utilizing the ability of energy-based model (EBM) to model complex distributions, EmSHAP provides an effective solution for estimating the required conditional probabilities. To further improve estimation accuracy, a GRU (Gated Recurrent Unit)-coupled partition function estimation method is introduced. The GRU network captures long-term dependencies with a lightweight parameterization and maps input features into a latent space to mitigate the influence of feature ordering. Additionally, a dynamic masking mechanism is incorporated to further enhance the robustness and accuracy by progressively increasing the masking rate. Theoretical analysis on the error bound as well as application to four case studies verified the higher accuracy and better scalability of EmSHAP in contrast to competitive methods.

cs.LG

Structured Sparsity Modeling for Improved Multivariate Statistical Analysis based Fault Isolation

In order to improve the fault diagnosis capability of multivariate statistical methods, this article introduces a fault isolation framework based on structured sparsity modeling. The developed method relies on the reconstruction based contribution analysis and the process structure information can be incorporated into the reconstruction objective function in the form of structured sparsity regularization terms. The structured sparsity terms allow selection of fault variables over structures like blocks or networks of process variables, hence more accurate fault isolation can be achieved. Four structured sparsity terms corresponding to different kinds of process information are considered, namely, partially known sparse support, block sparsity, clustered sparsity and tree-structured sparsity. The optimization problems involving the structured sparsity terms can be solved using the Alternating Direction Method of Multipliers (ADMM) algorithm, which is fast and efficient. Through a simulation example and an application study to a coal-fired power plant, it is verified that the proposed method can better isolate faulty variables by incorporating process structure information.

stat.AP

The limit of finite sample breakdown point of Tukey's halfspace median for general data

Under special conditions on data set and underlying distribution, the limit of finite sample breakdown point of Tukey's halfspace median ($\frac{1} {3}$) has been obtained in literature. In this paper, we establish the result under \emph{weaker assumption} imposed on underlying distribution (halfspace symmetry) and on data set (not necessary in general position). The representation of Tukey's sample depth regions for data set \emph{not necessary in general position} is also obtained, as a by-product of our derivation.

math.ST

Some results on the computing of Tukey's halfspace medain

Depth of the Tukey median is investigated for empirical distributions. A sharper upper bound is provided for this value for data sets in general position. This bound is lower than the existing one in the literature, and more importantly derived under the \emph{fixed} sample size practical scenario. Several results obtained in this paper are interesting theoretically and useful as well to reduce the computational burden of the Tukey median practically when $p$ is large relative to large $n$.

math.ST