SearcharxivSearch

arXiv subjects

Zongyue Teng

Publications and source records attributed to Zongyue Teng.

3 recordsLinked to original sources

Bayesian Variable Selection for High-Dimensional Predictors with Missing Psychometric Outcomes

High-dimensional, multimodal predictors and partially observed multivariate outcomes are common in psychometric research. However, existing regularization methods often do not accommodate hierarchical predictor structures and are primarily designed for univariate outcomes. We propose SHIM, a Bayesian framework for structured variable selection that combines hierarchical horseshoe shrinkage with a Bayesian treatment of missing outcomes. The framework jointly accommodates predictor hierarchies, dependence among outcomes, and incomplete multivariate responses. We establish theoretical properties of the proposed prior specification and evaluate SHIM through simulation studies. The results demonstrate that SHIM balances sensitivity with false-positive control while yielding accurate coefficient estimates and well-calibrated uncertainty quantification. We further apply SHIM to data from an Alzheimer's disease cohort to characterize associations between multimodal neuroimaging measures and multivariate neuropsychological outcomes and to generate posterior-based multiple imputations for downstream analyses of the relationships between fluid biomarkers and cognition. An R package, shim, is publicly available to facilitate implementation.

stat.ME

Nonlinear trajectories of lung function recovery in patients with pulmonary disease: empirical evaluation of longitudinal modeling approaches

Introduction: Longitudinal lung function recovery after pulmonary disease commonly follows nonlinear trajectories, and failure to adequately model these trajectories can lead to biased or misleading estimates of treatment effects. However, an important methodological gap remains as there is limited assessment of statistical methods for modeling nonlinear lung function trajectories. Methods: We compared several longitudinal modeling approaches for characterizing recovery in percent predicted forced expiratory volume in one second (FEV1p) using data from a phase 2 randomized trial for pulmonary tuberculosis (TB). We estimate the differences in repeated mean FEV1p between each treatment arm and control arm over a 180-days follow-up period. We compared 6 different statistical models: (1) linear mixed-effects model, (2) a piecewise linear mixed-effects model, (3) quadratic and (4) natural cubic spline mixed-effects models, (5) a nonlinear mixed-effects model with exponential recovery function, and (6) a generalized additive mixed model. We discussed the assumptions, clinical interpretations, and resulting treatment-effect estimates across these approaches. Results: The results from the TB trial analyses showed that the conventional linear mixed-effects model provided limited evidence of treatment differences over follow-up, whereas several flexible models identified significant differences during specific periods of recovery. Conclusion: Flexible longitudinal models can complement conventional linear mixed-effects models by revealing treatment differences at certain periods of follow-up that may be obscured by assuming a single linear trend. The choice of nonlinear modeling strategy should be guided by the scientific objective, available data, and the desired balance between clinical interpretability and flexibility.

stat.AP

When Does the Silhouette Score Work? A Comprehensive Study in Network Clustering

Selecting the number of communities is a fundamental challenge in network clustering. The silhouette score offers an intuitive, model-free criterion that balances within-cluster cohesion and between-cluster separation. Albeit its widespread use in clustering analysis, its performance in network-based community detection remains insufficiently characterized. In this study, we comprehensively evaluate the performance of the silhouette score across unweighted, weighted, and fully connected networks, examining how network size, separation strength, and community size imbalance influence its performance. Simulation studies show that the silhouette score accurately identifies the true number of communities when clusters are well separated and balanced, but it tends to underestimate under strong imbalance or weak separation and to overestimate in sparse networks. Extending the evaluation to a real airline reachability network, we demonstrate that the silhouette-based clustering can recover geographically interpretable and market-oriented clusters. These findings provide empirical guidance for applying the silhouette score in network clustering and clarify the conditions under which its use is most reliable.

cs.SI