SearcharxivSearch

arXiv subjects

Fuzhi Xu

Publications and source records attributed to Fuzhi Xu.

2 recordsLinked to original sources

Prediction-Powered Linear Regression: A Balance Between Interpretation and Prediction

Unlabeled data are increasingly prevalent in contemporary economic studies, yet their effective use for improving prediction remains challenging because the outcomes are often costly or even infeasible to observe. Machine learning methods can help label these data and achieve high predictive accuracy, but they often lack interpretability. In this paper, we propose a Prediction-powered Unified Model Averaging (PUMA) framework to combine linear regression and machine learning methods, achieving a balance between interpretation and prediction. Unlike existing studies on prediction-powered inference, our approach is the first to jointly address uncertainty arising from model misspecification, power tuning parameter selection, and the choice of machine learning algorithms by using model averaging. Theoretically, under mild conditions, we establish the in-sample and out-of-sample asymptotic prediction optimality, estimation consistency, and asymptotic distribution of the PUMA estimator. Extensive simulations and a real-world application further demonstrate the empirical advantages of the proposed method over existing state-of-the-art approaches.

stat.ME

Adaptive Multi-Prior Lasso for High-Dimensional Generalized Linear Models

Incorporation of external information into high-dimensional modeling for gene expression data has been shown, both theoretically and empirically, to substantially enhance performance. Such external information, sometimes referred to as prior information or priors, has become increasingly accessible from multiple sources, yet its reliability may vary considerably. Existing approaches often integrate these priors without sufficiently accounting for their quality, which may result in unsatisfactory or even misleading results. To effectively and selectively exploit such priors, we propose adaptive Multi-Prior Lasso, a novel regularization approach that simultaneously identifies reliable prior sources and integrates them to improve model performance. For high-dimensional generalized linear models (GLMs), an adaptive data-driven weight is assigned to each prior, so that more reliable sources are emphasized while less credible ones are downweighted. Theoretical guarantees are established, and the proposed method is shown through extensive simulations to improve estimation, prediction, and variable selection. An application to TCGA breast cancer gene expression data further illustrates the practical value of the proposed method, showing that incorporating prior information from PubMed published studies improves model performance.

stat.ME