SearcharxivSearch

arXiv subjects

Hwiyoung Lee

Publications and source records attributed to Hwiyoung Lee.

9 recordsLinked to original sources

Outcome-Calibrated Regression and Predicted Outcome-Based Inference

Regression is a fundamental tool in scientific research. Ordinary least squares (OLS), one of the most widely used regression methods, enjoys several desirable properties, including the best linear unbiased estimator (BLUE) property. It is well known that, under the assumptions of the standard model, the OLS is conditionally unbiased given the covariates, i.e., $\mathbb{E}(\widehat Y-Y\mid X=x)=0$. However, an often-overlooked property of OLS is that the prediction error is generally not unbiased conditional on the outcome, i.e., $\mathbb{E}(\widehat Y-Y\mid Y=y)\neq 0$. As a consequence of minimizing mean squared error, OLS predictions are systematically shrunk toward the outcome mean, which explains the classical phenomenon of regression to the mean (RTM): large outcome values tend to be underpredicted, whereas small outcome values tend to be overpredicted. This conditional prediction bias creates a nonignorable problem for predicted outcome-based inference, where scientific inference is performed using the predicted outcome $\widehat Y$ and another variable $W$. In applications such as brain-age analysis and causal inference, we show that inference based on regression-predicted outcomes can be systematically biased. To address this issue, we propose outcome-calibrated regression (OCR), a new regression framework with a closed-form solution that directly enforces outcome calibration. The proposed OCR estimator eliminates conditional prediction bias with respect to the outcome and enables valid inference using regression-predicted outcomes.

stat.ME

Trustworthy AI/ML Regression and Unbiased Causal Inference for Real-World Data

Real-World Data (RWD), with its large sample sizes and rich clinical detail, offers a compelling alternative to randomized controlled trials (RCTs) for studying treatment effects in diverse and complex patient populations. However, its observational nature introduces confounding that prevents straightforward comparative effectiveness research. Target trial emulation leverages RWD to estimate average treatment effects (ATE) at the population scale and diversity that RCTs cannot achieve, yet its validity depends critically on unbiased ATE estimation under high-dimensional confounding. Many causal inference pipelines address high-dimensional confounding through machine learning and artificial intelligence (ML/AI) outcome regression. However, commonly used ML/AI regression models exhibit systematic prediction bias, with predicted outcomes shrinking toward the marginal outcome mean. This structural bias propagates into ATE estimation and cannot be corrected by cross-fitting, ensemble methods, or any standard ML practice. In this work, we first quantitatively characterize how systematic prediction bias in ML/AI outcome regression leads to biased ATE estimates in causal inference models. We further propose an unbiased ML/AI regression-based causal inference framework to ensure unbiased ATE estimation for observational studies. We demonstrate our approach by studying the effects of opioids on cardiovascular health in patients with chronic pain using UK Biobank data.

stat.AP

Modeling Dependence in Omics Association Analysis via Structured Co-Expression Networks to Improve Power and Replicability

Accounting for dependence among high-dimensional variables in omics data analysis is critical to obtain accurate and reliable statistical inference. Although latent, omics variables often exhibit structured correlation/co-expression patterns. However, there are few methods explicitly accounting for such structured dependence in the statistical analysis of omics data (e.g., differential expression analysis). To address this methodological gap, we propose a Co-expression network multivariate Regression (CoReg), which integrates co-expression network structure into multivariate regression analysis to precisely account for the inter-correlations (dependence) among omics variables. We show in simulations that CoReg substantially improves the accuracy of statistical inference and replicability across studies. These findings suggest that CoReg provides an alternative approach for omics data association analysis with dependence adjustment, analogous to the role of mixed-effects models in handling repeated measures in lower-dimensional settings.

stat.ME

Graph Canonical Correlation Analysis

Canonical correlation analysis (CCA) is a widely used technique for estimating associations between two sets of multi-dimensional variables. Recent advancements in CCA methods have expanded their application to decipher the interactions of multiomics datasets, imaging-omics datasets, and more. However, conventional CCA methods are limited in their ability to incorporate structured patterns in the cross-correlation matrix, potentially leading to suboptimal estimations. To address this limitation, we propose the graph Canonical Correlation Analysis (gCCA) approach, which calculates canonical correlations based on the graph structure of the cross-correlation matrix between the two sets of variables. We develop computationally efficient algorithms for gCCA, and provide theoretical results for finite sample analysis of best subset selection and canonical correlation estimation by introducing concentration inequalities and stopping time rule based on martingale theories. Extensive simulations demonstrate that gCCA outperforms competing CCA methods. Additionally, we apply gCCA to a multiomics dataset of DNA methylation and RNA-seq transcriptomics, identifying both positively and negatively regulated gene expression pathways by DNA methylation pathways.

stat.ML

A Systematic Bias of Machine Learning Regression Models and Its Correction: an Application to Imaging-based Brain Age Prediction

Machine learning models for continuous outcomes often yield systematically biased predictions, particularly for values that largely deviate from the mean. Specifically, predictions for large-valued outcomes tend to be negatively biased (underestimating actual values), while those for small-valued outcomes are positively biased (overestimating actual values). We refer to this linear central tendency warped bias as the "systematic bias of machine learning regression". In this paper, we first demonstrate that this systematic prediction bias persists across various machine learning regression models, and then delve into its theoretical underpinnings. To address this issue, we propose a general constrained optimization approach designed to correct this bias and develop computationally efficient implementation algorithms. Simulation results indicate that our correction method effectively eliminates the bias from the predicted outcomes. We apply the proposed approach to the prediction of brain age using neuroimaging data. In comparison to competing machine learning regression models, our method effectively addresses the longstanding issue of "systematic bias of machine learning regression" in neuroimaging-based brain age calculation, yielding unbiased predictions of brain age.

stat.ML

Modeling Interconnected Modules in Multivariate Outcomes: Evaluating the Impact of Alcohol Intake on Plasma Metabolomics

Alcohol consumption has been shown to influence cardiovascular mechanisms in humans, leading to observable alterations in the plasma metabolomic profile. Regression models are commonly employed to investigate these effects, treating metabolomics features as the outcomes and alcohol intake as the exposure. Given the latent dependence structure among the numerous metabolomic features (e.g., co-expression networks with interconnected modules), modeling this structure is crucial for accurately identifying metabolomic features associated with alcohol intake. However, integrating dependence structures into regression models remains difficult in both estimation and inference procedures due to their large or high dimensionality. To bridge this gap, we propose an innovative multivariate regression model that accounts for correlations among outcome features by incorporating an interconnected community structure. Furthermore, we derive closed-form and likelihood-based estimators, accompanied by explicit exact and explicit asymptotic covariance matrix estimators, respectively. Simulation analysis demonstrates that our approach provides accurate estimation of both dependence and regression coefficients, and enhances sensitivity while maintaining a well-controlled discovery rate, as evidenced through benchmarking against existing regression models. Finally, we apply our approach to assess the impact of alcohol intake on $249$ metabolomic biomarkers measured using nuclear magnetic resonance spectroscopy. The results indicate that alcohol intake can elevate high-density lipoprotein levels by enhancing the transport rate of Apolipoproteins A1.

stat.ME

Robust Extrinsic Regression Analysis for Manifold Valued Data

Recently, there has been a growing need in analyzing data on manifolds owing to their important role in diverse fields of science and engineering. In the literature of manifold-valued data analysis up till now, however, only a few works have been carried out concerning the robustness of estimation against noises, outliers, and other sources of perturbations. In this regard, we introduce a novel extrinsic framework for analyzing manifold valued data in a robust manner. First, by extending the notion of the geometric median, we propose a new robust location parameter on manifolds, so-called the extrinsic median. A robust extrinsic regression method is also developed by incorporating the conditional extrinsic median into the classical local polynomial regression method. We present the Weiszfeld's algorithm for implementing the proposed methods. The promising performance of our approach against existing methods is illustrated through simulation studies.

stat.ME

Extrinsic Kernel Ridge Regression Classifier for Planar Kendall Shape Space

Kernel methods have had great success in Statistics and Machine Learning. Despite their growing popularity, however, less effort has been drawn towards developing kernel based classification methods on Riemannian manifolds due to difficulty in dealing with non-Euclidean geometry. In this paper, motivated by the extrinsic framework of manifold-valued data analysis, we propose a new positive definite kernel on planar Kendall shape space $Σ_2^k$, called extrinsic Veronese Whitney Gaussian kernel. We show that our approach can be extended to develop Gaussian kernels on any embedded manifold. Furthermore, kernel ridge regression classifier (KRRC) is implemented to address the shape classification problem on $Σ_2^k$, and their promising performances are illustrated through the real data analysis.

stat.ML

Anti-MANOVA on Compact Manifolds with Applications to 3D Projective Shape Analysis

Methods of hypotheses testing for equality of extrinsic antimeans on compact manifolds are unveiled in this paper. The two and multiple sample problem for antimeans on compact manifolds is addressed for large samples via asymptotic distributions, as well as for small samples using nonparametric bootstrap. An example of face differentiation using 3D VW antimean projective shape analysis for data extracted from digital camera images is also given.

math.ST