SearcharxivSearch

arXiv subjects

Xinzhou Guo

Publications and source records attributed to Xinzhou Guo.

14 recordsLinked to original sources

Optimal Inference of the Mean Outcome under Optimal Treatment Regime

When an optimal treatment regime (OTR) is considered, we need to evaluate the OTR in a valid and efficient way. The classical inference applied to the mean outcome under OTR, assuming the OTR is the same as the estimated OTR, might be biased when the regularity assumption that OTR is unique is violated. Although several methods have been proposed to allow nonregularity in such inference, its optimality is unclear due to challenges in deriving semiparametric efficiency bounds under potential nonregularity. In this paper, we address the bias issue via adaptive smoothing over the estimated OTR and develop a valid inference procedure on the mean outcome under OTR regardless of whether regularity is satisfied. We establish the optimality of the proposed method by deriving a lower bound of the asymptotic variance for the robust asymptotically linear unbiased estimator to the mean outcome under OTR and showing that our proposed estimator achieves the variance lower bound. The considered estimator class is general and the derived variance lower bound paves a novel way to establish efficiency optimality theories for OTR in a more general scenario allowing nonregularity. The merit of the proposed method is demonstrated by re-analyzing the ACTG 175 trial.

stat.ME

Efficiency Optimality without Pathwise Differentiability: A Variational Theory for Marginal-Integral Functionals

In this work, we provide a new perspective on semiparametric efficiency theory. In particular, we reformulate the questions of optimal efficiency and its attainment as a variational problem: minimize variance over estimating functions subject to robust unbiasedness constraints. We develop this theory for marginal-integral functionals without requiring pathwise differentiability. We study estimating functions whose expectations remain equal to the target when any one specified nuisance component is misspecified and the others are correct. We characterize the infimum of their variances through conditional variance minimization. For maxima of affine functions of treatment-specific conditional means, we obtain explicit optimal weights and construct cross-fitted estimators that attain the bound under conditions on nuisance estimation and the probability of near ties. For analogous maxima based on jointly observed quantities, the optimal weights use the full conditional covariance matrix. We also identify conditions under which the variance bound agrees with a classical convolution bound for parametric perturbations that preserve ties to first order. Examples include optimal policy values, $L^1$ calibration error, Balke--Pearl bounds, and mediation parameters. In an application to the National Longitudinal Survey of Young Men, covariance weighting reduces the median estimated variance of the cross-fitted estimating function relative to equal weighting by 12.9% for the lower Balke--Pearl endpoint and 16.5% for the upper, with improvements in all 20 repeated cross-fitting splits.

math.ST

Structure-based Transfer Learning

Transfer learning improves estimation in a target study using information from related sources. Classical transfer learning is typically data-based, requiring access to the source data or to a model fitted on them. Neither is available in many modern studies, as both are often proprietary or unreported. What can be transferred instead is structure: summarized information derived from a source, such as the supports of predictors, which is available and interpretable without access to the source data or model. The prevalence of the Large Language Model (LLM) provides a rich source of such structure information. However, data-based transfer learning can not well utilize it as accessing the data and model of the LLM is often infeasible. In this paper, we propose a novel transfer learning framework based on structure. In particular, we focus on linear model and aim to use supports generated by multiple agents, such as LLM, to improve estimation in the target study. Our procedure maps each support into the target parameter space, aggregates the results with data-driven weights, and applies a sparse correction against inaccurate supports. We establish nonasymptotic guarantees: the estimator improves on the target-only Lasso when the supports are informative and never does worse otherwise. Simulations and a gene-expression application confirm both properties.

stat.ME

Evaluation of Combination Therapy amid Patient-Level Heterogeneity

Combination therapy, a treatment approach that involves two or more monotherapies, is widely considered to enhance therapeutic efficacy across different medical conditions. It was often believed that combination therapy is beneficial because of pharmacological interactions between its component monotherapies. However, through laboratory experiments, pharmacologists have recently noted that the benefits of some combination therapies might be largely driven by varying patient-level responses to their component monotherapies. Without accounting for such patient-level heterogeneity, classical statistical inference frameworks for combination therapy might be inadequate and overly optimistic. In this paper, we introduce a novel and model-free statistical inference framework to complement the classical one and evaluate combination therapy after adjusting for patient-level heterogeneity in responses to monotherapies. We address the non-identifiability and nonlinearity issue inherent in adjustment of patient-level heterogeneity and establish conditions for the (partial) identifiability of the cross-world target parameter. We develop an outcome-based optimal matching scheme to achieve asymptotic normality and construct $\sqrt{N}$-rate confidence intervals for the target parameter, thereby enabling reliable, efficient and transparent evaluation of combination therapy amid patient-level heterogeneity. The benefits of the proposed framework are demonstrated through a reanalysis of the ACTG 175 trial.

stat.ME

Stein-Encoder: A White-Box Supervised Encoder via Stein Identities in Multi-Modal Studies

In multi-modal biomedical research, integrating high-dimensional genomic data with clinical baselines is essential for precision medicine. However, standard deep neural network approaches often entangle these modalities, obscuring the specific predictive impact of genetic features and leading to possibly suboptimal predictive performance. Motivated by the landmark METABRIC cohort primary breast tumors study, we propose the Stein-Encoder, a white-box supervised framework designed to isolate the genetic signal driving clinical outcomes conditional on nuisance covariates. By leveraging Stein's method and residualization techniques, our approach constructs an interpretable single index that summarizes relevant biological heterogeneity while flexibly incorporating clinical factors and can be used to improve downstream prediction. We establish theoretical guarantees for identification, consistency and efficiency improvement. Applied to the METABRIC cohort, the Stein-Encoder outperforms unsupervised benchmarks in predictive accuracy. Crucially, it achieves structural disentanglement by revealing response-specific biological mechanisms: we find that tumor size is driven primarily by mitotic networks, whereas prognostic indices rely on a distinct proliferation-versus-immune axis. This work contributes a unified, computationally efficient framework that bridges statistical rigor with the representational power of neural networks, enabling interpretable, task-specific and efficient compression of multi-modal health data for a wide range of precision medicine applications, beyond biomarker discovery.

stat.AP

In-Sample Evaluation of Subgroups Identified by Generic Machine Learning

When a subgroup is identified from the data, it must be evaluated in a replicable way. The usual in-sample approach, which evaluates the post-hoc identified subgroup as predefined, might suffer from selection bias. This issue of in-sample evaluation of data-dependent objects is well recognized but particularly challenging here. Unlike discrete or finite-dimensional data-dependent objects addressed before, the selection bias here is induced by post-hoc identified subgroups, data-dependent sets potentially defined by infinite-dimensional functionals with nonsmooth boundaries known as nonregularity. The out-of-sample approach, which splits data for subgroup identification and evaluation, can help address selection bias but might suffer from efficiency loss and instability. In this paper, we propose a conditional adaptive perturbation approach to remove selection bias in in-sample subgroup evaluation and deliver valid inference on subgroups identified from the whole dataset by generic machine learning, regardless of whether regularity is satisfied. The proposed method is easy-to-compute, allows model-free and even black-box subgroup identification, and achieves full efficiency across broad scenarios of subgroup analysis through a novel theoretical framework of triple robustness linking rates of subgroup identification and nuisance estimation. The merits of the proposed method are demonstrated by a re-analysis of the ACTG 175 trial.

stat.ME

Valid and Efficient Two-Stage Latent Subgroup Analysis with Observational Data

Subgroup analysis evaluates treatment effects across multiple sub-populations. When subgroups are defined by latent memberships inferred from imperfect measurements, the analysis typically involves two inter-connected models, a latent class model and a subgroup outcome model. The classical one-stage framework, which models the joint distribution of the two models, may be infeasible with observational data containing many confounders. The two-stage framework, which first estimates the latent class model and then performs subgroup analysis using estimated latent memberships, can accommodate potential confounders but may suffer from bias issues due to misclassification of latent subgroup memberships. This paper focuses on latent subgroups inferred from binary item responses and addresses when and how a valid two-stage latent subgroup analysis can be made with observational data. We investigate the maximum misclassification rate that a valid two-stage framework can tolerate. Introducing a spectral method perspective, we propose a two-stage approach to achieve the desired misclassification rate with the blessing of many item responses. Our method accommodates high-dimensional confounders, is computationally efficient and robust to noninformative items. In observational studies, our methods lead to consistent estimation and valid inference on latent subgroup effects. We demonstrate its merit through simulation studies and an application to educational assessment data.

stat.ME

Study on the Impact of Radioactive Background on the Dark Count Rate of 20-inch MCP-PMTs

This study systematically investigates the impact of natural radioactive background on the dark count rate (DCR) of 20-inch microchannel plate photomultiplier tubes (MCP-PMTs). Variations in PMT DCR under different radiation conditions were examined via underground tests, lead shielding experiments, and irradiation with \(^{55}\mathrm{Fe}\), \(^{60}\mathrm{Co}\), and \(^{90}\mathrm{Sr}\) sources. The experimental results indicate that natural radioactivity from surrounding air and rock in the underground environment results in a significantly higher DCR compared to laboratory conditions. Further, lead shielding experiments confirmed that effective shielding can markedly reduce the interference from environmental background radiation. Notably, $β$ particles from the \(^{90}\mathrm{Sr}\) source increased the DCR by approximately 14 kHz, whereas the effects of $\mathrm{X}$-rays from \(^{55}\mathrm{Fe}\) and $γ$-rays from \(^{60}\mathrm{Co}\) were comparatively minor. In addition, Geant4 simulations provided quantitative analysis of Cherenkov radiation induced by $β$ particles, with the simulation results closely matching the experimental data and confirming $β$ particles as the primary contributor to the DCR increase. These findings offer both theoretical and experimental evidence for a deeper understanding of the influence of radioactive background on 20-inch MCP-PMTs' performance in underground experiments and hold significant implications for improving the energy resolution of large-scale neutrino detection systems

hep-ex

Bootstrapping the Cross-Validation Estimate

Cross-validation is a widely used technique for evaluating the performance of prediction models, ranging from simple binary classification to complex precision medicine strategies. It helps correct for optimism bias in error estimates, which can be significant for models built using complex statistical learning algorithms. However, since the cross-validation estimate is a random value dependent on observed data, it is essential to accurately quantify the uncertainty associated with the estimate. This is especially important when comparing the performance of two models using cross-validation, as one must determine whether differences in estimated error are due to chance. Although various methods have been developed to make inferences on cross-validation estimates, they often have many limitations, such as requiring stringent model assumptions. This paper proposes a fast bootstrap method that quickly estimates the standard error of the cross-validation estimate and produces valid confidence intervals for a population parameter measuring average model performance. Our method overcomes the computational challenges inherent in bootstrapping a cross-validation estimate by estimating the variance component within a random-effects model. It is also as flexible as the cross-validation procedure itself. To showcase the effectiveness of our approach, we conducted comprehensive simulations and real-data analysis across two applications.

stat.ME

A Test System for the JUNO 20-inch PMTs Prior to Installation

The JUNO experiment requires an excellent energy resolution of 3\% at 1 MeV. To achieve this objective, a total of 20,012 20-inch photomultiplier tubes (PMTs) will be deployed for JUNO, comprising 15,012 multi-channel plate (MCP) PMTs and 5,000 dynode PMTs. Currently, JUNO is in the process of detector installation, with PMTs being installed from the top to the bottom of the stainless-steel structure located in the underground experimental hall. In order to validate the functionality of the PMTs and ensure there are no malfunctions prior to installation, a test system has been established at the JUNO site, and testing is being conducted. This paper presents an overview of the test system and reports on the initial test results.

physics.ins-det

Differentially Private Multimodal Laplacian Dropout (DP-MLD) for EEG Representative Learning

Recently, multimodal electroencephalogram (EEG) learning has shown great promise in disease detection. At the same time, ensuring privacy in clinical studies has become increasingly crucial due to legal and ethical concerns. One widely adopted scheme for privacy protection is differential privacy (DP) because of its clear interpretation and ease of implementation. Although numerous methods have been proposed under DP, it has not been extensively studied for multimodal EEG data due to the complexities of models and signal data considered there. In this paper, we propose a novel Differentially Private Multimodal Laplacian Dropout (DP-MLD) scheme for multimodal EEG learning. Our approach proposes a novel multimodal representative learning model that processes EEG data by language models as text and other modal data by vision transformers as images, incorporating well-designed cross-attention mechanisms to effectively extract and integrate cross-modal features. To achieve DP, we design a novel adaptive feature-level Laplacian dropout scheme, where randomness allocation and performance are dynamically optimized within given privacy budgets. In the experiment on an open-source multimodal dataset of Freezing of Gait (FoG) in Parkinson's Disease (PD), our proposed method demonstrates an approximate 4\% improvement in classification accuracy, and achieves state-of-the-art performance in multimodal EEG learning under DP.

eess.SP

Differentially Private Confidence Interval for Extrema of Parameters

This paper aims to construct a valid and efficient confidence interval for the extrema of parameters under privacy protection. The usual statistical inference on the extrema of parameters often suffers from the selection bias issue, and the problem becomes more acute, as in many application scenarios of extrema parameters, we often need to protect the privacy of the data. In this paper, we focus on the exponential family of distributions and propose a privatized parametric bootstrap method to address selection bias in the extrema of parameters problem under the scheme of differential privacy. While the usual privatized parametric bootstrap does not address selection bias appropriately, we prove that with a privatized bias correction term, the proposed parametric bootstrap method can lead to a valid and efficient confidence interval for the extrema of parameters. We illustrate the proposed method with the Gaussian case and regression case and demonstrate the advantages of the proposed method via numerical experiments.

stat.ME

Assessing the Most Vulnerable Subgroup to Type II Diabetes Associated with Statin Usage: Evidence from Electronic Health Record Data

There have been increased concerns that the use of statins, one of the most commonly prescribed drugs for treating coronary artery disease, is potentially associated with the increased risk of new-onset type II diabetes (T2D). Nevertheless, to date, there is no robust evidence supporting as to whether and what kind of populations are indeed vulnerable for developing T2D after taking statins. In this case study, leveraging the biobank and electronic health record data in the Partner Health System, we introduce a new data analysis pipeline and a novel statistical methodology that address existing limitations by (i) designing a rigorous causal framework that systematically examines the causal effects of statin usage on T2D risk in observational data, (ii) uncovering which patient subgroup is most vulnerable for developing T2D after taking statins, and (iii) assessing the replicability and statistical significance of the most vulnerable subgroup via a bootstrap calibration procedure. Our proposed approach delivers asymptotically sharp confidence intervals and debiased estimate for the treatment effect of the most vulnerable subgroup in the presence of high-dimensional covariates. With our proposed approach, we find that females with high T2D genetic risk are at the highest risk of developing T2D due to statin usage.

stat.AP

Sharp Inference on Selected Subgroups in Observational Studies

In modern drug development, the broader availability of high-dimensional observational data provides opportunities for scientist to explore subgroup heterogeneity, especially when randomized clinical trials are unavailable due to cost and ethical constraints. However, a common practice that naively searches the subgroup with a high treatment level is often misleading due to the "subgroup selection bias." More importantly, the nature of high-dimensional observational data has further exacerbated the challenge of accurately estimating the subgroup treatment effects. To resolve these issues, we provide new inferential tools based on resampling to assess the replicability of post-hoc identified subgroups from observational studies. Through careful theoretical justification and extensive simulations, we show that our proposed approach delivers asymptotically sharp confidence intervals and debiased estimates for the selected subgroup treatment effects in the presence of high-dimensional covariates. We further demonstrate the merit of the proposed methods by analyzing the UK Biobank data.

stat.ME