SearcharxivSearch

arXiv subjects

Shuoxun Xu

Publications and source records attributed to Shuoxun Xu.

7 recordsLinked to original sources

Structure-based Transfer Learning

Transfer learning improves estimation in a target study using information from related sources. Classical transfer learning is typically data-based, requiring access to the source data or to a model fitted on them. Neither is available in many modern studies, as both are often proprietary or unreported. What can be transferred instead is structure: summarized information derived from a source, such as the supports of predictors, which is available and interpretable without access to the source data or model. The prevalence of the Large Language Model (LLM) provides a rich source of such structure information. However, data-based transfer learning can not well utilize it as accessing the data and model of the LLM is often infeasible. In this paper, we propose a novel transfer learning framework based on structure. In particular, we focus on linear model and aim to use supports generated by multiple agents, such as LLM, to improve estimation in the target study. Our procedure maps each support into the target parameter space, aggregates the results with data-driven weights, and applies a sparse correction against inaccurate supports. We establish nonasymptotic guarantees: the estimator improves on the target-only Lasso when the supports are informative and never does worse otherwise. Simulations and a gene-expression application confirm both properties.

stat.ME

Evaluation of Combination Therapy amid Patient-Level Heterogeneity

Combination therapy, a treatment approach that involves two or more monotherapies, is widely considered to enhance therapeutic efficacy across different medical conditions. It was often believed that combination therapy is beneficial because of pharmacological interactions between its component monotherapies. However, through laboratory experiments, pharmacologists have recently noted that the benefits of some combination therapies might be largely driven by varying patient-level responses to their component monotherapies. Without accounting for such patient-level heterogeneity, classical statistical inference frameworks for combination therapy might be inadequate and overly optimistic. In this paper, we introduce a novel and model-free statistical inference framework to complement the classical one and evaluate combination therapy after adjusting for patient-level heterogeneity in responses to monotherapies. We address the non-identifiability and nonlinearity issue inherent in adjustment of patient-level heterogeneity and establish conditions for the (partial) identifiability of the cross-world target parameter. We develop an outcome-based optimal matching scheme to achieve asymptotic normality and construct $\sqrt{N}$-rate confidence intervals for the target parameter, thereby enabling reliable, efficient and transparent evaluation of combination therapy amid patient-level heterogeneity. The benefits of the proposed framework are demonstrated through a reanalysis of the ACTG 175 trial.

stat.ME

Stein-Encoder: A White-Box Supervised Encoder via Stein Identities in Multi-Modal Studies

In multi-modal biomedical research, integrating high-dimensional genomic data with clinical baselines is essential for precision medicine. However, standard deep neural network approaches often entangle these modalities, obscuring the specific predictive impact of genetic features and leading to possibly suboptimal predictive performance. Motivated by the landmark METABRIC cohort primary breast tumors study, we propose the Stein-Encoder, a white-box supervised framework designed to isolate the genetic signal driving clinical outcomes conditional on nuisance covariates. By leveraging Stein's method and residualization techniques, our approach constructs an interpretable single index that summarizes relevant biological heterogeneity while flexibly incorporating clinical factors and can be used to improve downstream prediction. We establish theoretical guarantees for identification, consistency and efficiency improvement. Applied to the METABRIC cohort, the Stein-Encoder outperforms unsupervised benchmarks in predictive accuracy. Crucially, it achieves structural disentanglement by revealing response-specific biological mechanisms: we find that tumor size is driven primarily by mitotic networks, whereas prognostic indices rely on a distinct proliferation-versus-immune axis. This work contributes a unified, computationally efficient framework that bridges statistical rigor with the representational power of neural networks, enabling interpretable, task-specific and efficient compression of multi-modal health data for a wide range of precision medicine applications, beyond biomarker discovery.

stat.AP

In-Sample Evaluation of Subgroups Identified by Generic Machine Learning

When a subgroup is identified from the data, it must be evaluated in a replicable way. The usual in-sample approach, which evaluates the post-hoc identified subgroup as predefined, might suffer from selection bias. This issue of in-sample evaluation of data-dependent objects is well recognized but particularly challenging here. Unlike discrete or finite-dimensional data-dependent objects addressed before, the selection bias here is induced by post-hoc identified subgroups, data-dependent sets potentially defined by infinite-dimensional functionals with nonsmooth boundaries known as nonregularity. The out-of-sample approach, which splits data for subgroup identification and evaluation, can help address selection bias but might suffer from efficiency loss and instability. In this paper, we propose a conditional adaptive perturbation approach to remove selection bias in in-sample subgroup evaluation and deliver valid inference on subgroups identified from the whole dataset by generic machine learning, regardless of whether regularity is satisfied. The proposed method is easy-to-compute, allows model-free and even black-box subgroup identification, and achieves full efficiency across broad scenarios of subgroup analysis through a novel theoretical framework of triple robustness linking rates of subgroup identification and nuisance estimation. The merits of the proposed method are demonstrated by a re-analysis of the ACTG 175 trial.

stat.ME

Robust Learning of Heterogeneous Dynamic Systems

Ordinary differential equations (ODEs) provide a powerful framework for modeling dynamic systems arising in a wide range of scientific domains. However, most existing ODE methods focus on a single system, and do not adequately address the problem of learning shared patterns from multiple heterogeneous dynamic systems. In this article, we propose a novel distributionally robust learning approach for modeling heterogeneous ODE systems. Specifically, we construct a robust dynamic system by maximizing a worst-case reward over an uncertainty class formed by convex combinations of the derivatives of trajectories. We show the resulting estimator admits an explicit weighted average representation, where the weights are obtained from a quadratic optimization that balances information across multiple data sources. We further develop a bi-level stabilization procedure to address potential instability in estimation. We establish rigorous theoretical guarantees for the proposed method, including consistency of the stabilized weights, error bound for robust trajectory estimation, and asymptotical validity of pointwise confidence interval. We demonstrate that the proposed method considerably improves the generalization performance compared to the alternative solutions through both extensive simulations and the analysis of an intracranial electroencephalogram data.

stat.ME

Statistical Learning for Latent Embedding Alignment with Application to Brain Encoding and Decoding

Brain encoding and decoding aims to understand the relationship between external stimuli and brain activities, and is a fundamental problem in neuroscience. In this article, we study latent embedding alignment for brain encoding and decoding, with a focus on improving sample efficiency under limited fMRI-stimulus paired data and substantial subject heterogeneity. We propose a lightweight alignment framework equipped with two statistical learning components: inverse semi-supervised learning that leverages abundant unpaired stimulus embeddings through inverse mapping and residual debiasing, and meta transfer learning that borrows strength from pretrained models across subjects via sparse aggregation and residual correction. Both methods operate exclusively at the alignment stage while keeping encoders and decoders frozen, allowing for efficient computation, modular deployment, and rigorous theoretical analysis. We establish finite-sample generalization bounds and safety guarantees, and demonstrate competitive empirical performance on the large-scale fMRI-image reconstruction benchmark data.

stat.ME

Optimal Inference of the Mean Outcome under Optimal Treatment Regime

When an optimal treatment regime (OTR) is considered, we need to evaluate the OTR in a valid and efficient way. The classical inference applied to the mean outcome under OTR, assuming the OTR is the same as the estimated OTR, might be biased when the regularity assumption that OTR is unique is violated. Although several methods have been proposed to allow nonregularity in such inference, its optimality is unclear due to challenges in deriving semiparametric efficiency bounds under potential nonregularity. In this paper, we address the bias issue via adaptive smoothing over the estimated OTR and develop a valid inference procedure on the mean outcome under OTR regardless of whether regularity is satisfied. We establish the optimality of the proposed method by deriving a lower bound of the asymptotic variance for the robust asymptotically linear unbiased estimator to the mean outcome under OTR and showing that our proposed estimator achieves the variance lower bound. The considered estimator class is general and the derived variance lower bound paves a novel way to establish efficiency optimality theories for OTR in a more general scenario allowing nonregularity. The merit of the proposed method is demonstrated by re-analyzing the ACTG 175 trial.

stat.ME