SearcharxivSearch

arXiv subjects

Judong Shen

Publications and source records attributed to Judong Shen.

5 recordsLinked to original sources

Uncertainty-Aware Multimodal Learning via Conformal Shapley Intervals

Multimodal learning combines information from multiple data modalities to improve predictive performance. However, modalities often contribute unequally and in a data dependent way, making it unclear which data modalities are genuinely informative and to what extent their contributions can be trusted. Quantifying modality level importance together with uncertainty is therefore central to interpretable and reliable multimodal learning. We introduce conformal Shapley intervals, a framework that combines Shapley values with conformal inference to construct uncertainty-aware importance intervals for each modality. Building on these intervals, we propose a modality selection procedure with a provable optimality guarantee: conditional on the observed features, the selected subset of modalities achieves performance close to that of the optimal subset. We demonstrate the effectiveness of our approach on multiple datasets, showing that it provides meaningful uncertainty quantification and strong predictive performance while relying on only a small number of informative modalities.

stat.ML

Estimating heritability of survival traits using censored multiple variance component model

Characterizing the genetic basis of survival traits, such as age at disease onset, is critical for risk stratification, early intervention, and elucidating biological mechanisms that can inform therapeutic development. However, time-to-event outcomes in human cohorts are frequently right-censored, complicating both the estimation and partitioning of total heritability. Modern biobanks linked to electronic health records offer the unprecedented power to dissect the genetic basis of age-at-diagnosis traits at large scale. Yet, few methods exist for estimating and partitioning the total heritability of censored survival traits. Existing methods impose restrictive distributional assumptions on genetic and environmental effects and are not scalable to large biobanks with a million subjects. We introduce a censored multiple variance component model to robustly estimate the total heritability of survival traits under right-censoring. We demonstrate through extensive simulations that the method provides accurate total heritability estimates of right-censored traits at censoring rates up to 80% given sufficient sample size. The method is computationally efficient in estimating one hundred genetic variance components of a survival trait using large-scale biobank genotype data consisting of a million subjects and a million SNPs in under nine hours, including uncertainty quantification. We apply our method to estimate the total heritability of four age-at-diagnosis traits from the UK Biobank study. Our results establish a scalable and robust framework for heritability analysis of right-censored survival traits in large-scale genetic studies.

stat.ME

CauchyCP: a powerful test under non-proportional hazards using Cauchy combination of change-point Cox regressions

Non-proportional hazards data are routinely encountered in randomized clinical trials. In such cases, classic Cox proportional hazards model can suffer from severe power loss, with difficulty in interpretation of the estimated hazard ratio since the treatment effect varies over time. We propose CauchyCP, an omnibus test of change-point Cox regression models, to overcome both challenges while detecting signals of non-proportional hazards patterns. Extensive simulation studies demonstrate that, compared to existing treatment comparison tests under non-proportional hazards, the proposed CauchyCP test 1) controls the type I error better at small $α$ levels ($< 0.01$); 2) increases the power of detecting time-varying effects; and 3) is more computationally efficient. The superior performance of CauchyCP is further illustrated using retrospective analyses of two randomized clinical trial datasets and a pharmacogenetic biomarker study dataset. The R package $\textit{CauchyCP}$ is publicly available on CRAN.

stat.ME

An efficient and accurate approximation to the distribution of quadratic forms of Gaussian variables

In computational and applied statistics, it is of great interest to get fast and accurate calculation for the distributions of the quadratic forms of Gaussian random variables. This paper presents a novel approximation strategy that contains two developments. First, we propose a faster numerical procedure in computing the moments of the quadratic forms. Second, we establish a general moment-matching framework for distribution approximation, which covers existing approximation methods for the distributions of the quadratic forms of Gaussian variables. Under this framework, a novel moment-ratio method (MR) is proposed to match the ratio of skewness and kurtosis based on the gamma distribution. Our extensive simulations show that 1) MR is almost as accurate as the exact distribution calculation and is much more efficient; 2) comparing with existing approximation methods, MR significantly improves the accuracy of approximating far right tail probabilities. The proposed method has wide applications. For example, it is a better choice than existing methods for facilitating hypothesis testing in big data analysis, where efficient and accurate calculation of very small $p$-values is desired. An R package Qapprox that implements related methods is available on CRAN.

stat.ME

A Random Interaction Forest for Prioritizing Predictive Biomarkers

Precision medicine is becoming a focus in medical research recently, as its implementation brings values to all stakeholders in the healthcare system. Various statistical methodologies have been developed tackling problems in different aspects of this field, e.g., assessing treatment heterogeneity, identifying patient subgroups, or building treatment decision models. However, there is a lack of new tools devoted to selecting and prioritizing predictive biomarkers. We propose a novel tree-based ensemble method, random interaction forest (RIF), to generate predictive importance scores and prioritize candidate biomarkers for constructing refined treatment decision models. RIF was evaluated by comparing with the conventional random forest and univariable regression methods and showed favorable properties under various simulation scenarios. We applied the proposed RIF method to a biomarker dataset from two phase III clinical trials of bezlotoxumab on $\textit{Clostridium difficile}$ infection recurrence and obtained biologically meaningful results.

q-bio.QM