SearcharxivSearch

arXiv subjects

Antonio Linero

Publications and source records attributed to Antonio Linero.

5 recordsLinked to original sources

Bayesian machine learning approach for recurrent events studies using Soft Bayesian Additive Regression Trees (SBART)

Recurrent event data frequently arise in biomedical studies, where individuals may experience multiple recurrences of the same type of events, such as recurrent hospitalizations. This article introduces a nonparametric method for recurrent events under a Bayesian ensemble learning framework, called Soft Bayesian Additive Regression Trees (SBART), which combines multiple soft decision trees to achieve high predictive accuracy and a smooth estimator of the underlying intensity of the recurrent events. The proposed model represents the conditional intensity function of the non-homogeneous Poisson process as the product of a time-constant baseline, a subject-specific frailty random effect, and a nonparametric component capturing potentially nonlinear covariate effects and unknown interactions among covariates and time. A two-layer data augmentation scheme is employed to efficiently incorporate the SBART component within our computational algorithm. Simulation studies demonstrate that our method, called RecSBART in short, achieves superior accuracy in estimating cumulative intensity compared to existing approaches, even when our modeling assumptions are not true. With the Bayesian analysis of a study of recurrent hospitalizations of colorectal cancer patients, we further demonstrate our RecSBART method's ability to reveal and interpret the underlying complex relationships among covariates in a recurrent events study.

stat.ME

Treatment Effect Heterogeneity and Importance Measures for Multivariate Continuous Treatments

Estimating the joint effect of a multivariate, continuous exposure is crucial, particularly in environmental health where interest lies in simultaneously evaluating the impact of multiple environmental pollutants on health. We develop novel methodology that addresses two key issues for estimation of treatment effects of multivariate, continuous exposures. We use nonparametric Bayesian methodology that is flexible to ensure our approach can capture a wide range of data generating processes. Additionally, we allow the effect of the exposures to be heterogeneous with respect to covariates. Treatment effect heterogeneity has not been well explored in the causal inference literature for multivariate, continuous exposures, and therefore we introduce novel estimands that summarize the nature and extent of the heterogeneity, and propose estimation procedures for new estimands related to treatment effect heterogeneity. We provide theoretical support for the proposed models in the form of posterior contraction rates and show that it works well in simulated examples both with and without heterogeneity. Our approach is motivated by a study of the health effects of simultaneous exposure to the components of PM$_{2.5}$, where we find that the negative health effects of exposure to environmental pollutants are exacerbated by low socioeconomic status, race and age.

stat.ME

Predictive variational inference: Learn the predictively optimal posterior distribution

Vanilla variational inference finds an optimal approximation to the Bayesian posterior distribution, but even the exact Bayesian posterior is often not meaningful under model misspecification. We propose predictive variational inference (PVI): a general inference framework that seeks and samples from an optimal posterior density such that the resulting posterior predictive distribution is as close to the true data generating process as possible, while this closeness is measured by multiple scoring rules. By optimizing the objective, the predictive variational inference is generally not the same as, or even attempting to approximate, the Bayesian posterior, even asymptotically. Rather, we interpret it as implicit hierarchical expansion. Further, the learned posterior uncertainty detects heterogeneity of parameters among the population, enabling automatic model diagnosis. This framework applies to both likelihood-exact and likelihood-free models. We demonstrate its application in real data examples.

stat.ML

Density Regression with Bayesian Additive Regression Trees

Flexibly modeling how an entire density changes with covariates is an important but challenging generalization of mean and quantile regression. While existing methods for density regression primarily consist of covariate-dependent discrete mixture models, we consider a continuous latent variable model in general covariate spaces, which we call DR-BART. The prior mapping the latent variable to the observed data is constructed via a novel application of Bayesian Additive Regression Trees (BART). We prove that the posterior induced by our model concentrates quickly around true generative functions that are sufficiently smooth. We also analyze the performance of DR-BART on a set of challenging simulated examples, where it outperforms various other methods for Bayesian density regression. Lastly, we apply DR-BART to two real datasets from educational testing and economics, to study student growth and predict returns to education. Our proposed sampler is efficient and allows one to take advantage of BART's flexibility in many applied settings where the entire distribution of the response is of primary interest. Furthermore, our scheme for splitting on latent variables within BART facilitates its future application to other classes of models that can be described via latent variables, such as those involving hierarchical or time series data.

stat.ME