SearcharxivSearch

arXiv subjects

Yangsheng Wang

Publications and source records attributed to Yangsheng Wang.

3 recordsLinked to original sources

Theoretical Properties of Multivariate Random Forest in Feature Selection and its Application to Facial Morphology-Gene Detection

This work establishes a theoretical foundation for joint feature selection with multivariate outcomes, positioning the permutation-based variable importance measure (PVIM) of multivariate random forests (MRF) as a principled tool for high-dimensional feature selection. We establish the first consistency guaranty for MRF, showing that it retains all truly influential features with probability tending to one as the sample size grows to infinity under mild regularity conditions. Incomplete U-statistics is employed to incorporate three layers of randomness: subsampling of subjects for training each tree, subsampling of features at each split, and permutation of each feature for the out-of-bag (OOB) samples. Unlike independence-based screening that evaluates each feature in isolation, PVIM is a joint screening approach that accounts for multicollinearity, nonlinear, high-order interactions, and subject heterogeneity via ensemble aggregation. Moreover, we demonstrate the practical utility of MRF through a genome-wide association study (GWAS) of human facial morphology (with 2,342 subjects and 453,273 SNPs), where MRF identifies several novel loci and interaction hubs that extend prior findings. Extensive simulations show that MRF accurately identifies truly influential signals while producing parsimonious feature sets with well controlled false selection rates, outperforming canonical correlation analysis (CCA) and several other independence multivariate screening approaches. In addition, we also propose a novel simulation framework, including image outcomes, that more closely mimic the intricate nature of real-world data and provide rigorous testbeds for machine learning research.

stat.ME

Longitudinal Random Forests for Sparse and Irregular Response Trajectories

Longitudinal studies often collect data at sparse, irregular, and unequally spaced time points. Such heterogeneity is often driven by subject-specific covariates, yet existing methods have been restricted to a scalar endpoint value, completely neglecting the underlying response trajectories. We propose a novel Longitudinal Random Forest (LRF) framework that leverages tree-based ensemble machine learning with adaptive node-wise longitudinal trajectory estimation. The LRF framework makes five methodological contributions. it captures each subject's individual response trajectory while simultaneously accommodating within-node correlation, between-node heterogeneity, and nonlinear and interactive covariate effects. It introduces a novel trajectory-based splitting criterion that maximizes trajectory separation while incorporating a size-weighted penalty; it provides two variants, Principal Analysis by Conditional Expectation (LRF-PACE) and adaptive linear mixed-effects models (LRF-adaptiveLMM), which employ nonparametric and semiparametric node-wise smoothers, respectively, while learning covariate effects in a data-driven manner. It provides a comprehensive interpretation of covariates using both the classical trajectory-based permutation variable importance measure (PVIM) and a newly proposed finite-way interaction frequency count, and it not only predicts entire trajectories for new subjects but also forecasts future trajectories for existing subjects. Extensive simulation studies demonstrate that LRF achieves superior performance over several competing methods, even under severe sparsity. The practical significance of the LRF framework lies in its ability to address five important clinical questions.

stat.ME

An Integrated Deep Learning and Statistical Framework for Whole-Network Gene--Environment Association with Leaf Vascular Architecture

Leaf veins exhibit remarkable diversity in architecture and patterning, yet existing gene--environment association studies have primarily quantified leaf venation using a small collection of low-dimensional summary traits, thereby discarding most of the structural information contained in the original images. We propose an integrated deep learning and statistical framework. The proposed framework achieves four methodological advances. First, it represents the complete leaf vascular architecture as a whole-network image phenotype. Second, it fine-tunes the deep learning-based Edge Detection with Transformers (EDTER) model to accurately extract whole-network leaf vascular architecture from RGB images by jointly learning local and global contextual features. Third, it constructs a new annotated leaf image database by integrating edge maps generated by DiffusionEdge with the Berkeley Segmentation Database (BSDS500). Fourth, it applies Semiparametric Sparse Canonical Correlation Analysis (SSCCA) to perform variable selection and model associations between repeatedly measured high-dimensional Bivariate image responses and high-dimensional predictors while simultaneously accommodating sparse, zero-inflated data represented by edge maps through a truncated latent Gaussian copula model. Two simulation studies demonstrate the performance of the proposed framework under increasing levels of complexity. Application to a real \emph{Populus} dataset identifies three significant gene--geography interactions associated with leaf vascular architecture, providing new biological insights and establishing a broadly applicable methodological framework for high-dimensional complex image phenotypes.

cs.LG