SearcharxivSearch

arXiv subjects

Daniel Kowal

Publications and source records attributed to Daniel Kowal.

3 recordsLinked to original sources

Bayesian Quantile Regression with Subset Selection: A Decision Analysis Perspective

Quantile regression is a powerful tool for inferring how covariates affect specific percentiles of the response distribution. Existing methods either estimate conditional quantiles separately for each quantile of interest or estimate the entire conditional distribution using semi- or non-parametric models. The former often produce inadequate models for real data and do not share information across quantiles, while the latter are characterized by complex and constrained models that can be difficult to interpret and computationally inefficient. Neither approach is well-suited for quantile-specific subset selection. Instead, we pose the fundamental problems of linear quantile estimation, uncertainty quantification, and subset selection from a Bayesian decision analysis perspective. For any Bayesian regression model -- including, but not limited to existing Bayesian quantile regression models -- we derive optimal point estimates, interpretable uncertainty quantification, and scalable subset selection techniques for all model-based conditional quantiles. Our approach introduces a quantile-focused squared error loss that enables efficient, closed-form computing and maintains a close relationship with Wasserstein-based density estimation. In an extensive simulation study, our methods demonstrate substantial gains in quantile estimation accuracy, inference, and variable selection over frequentist and Bayesian competitors. We use these tools to identify and quantify the heterogeneous impacts of multiple social stressors and environmental exposures on educational outcomes across the full spectrum of low-, medium-, and high-achieving students in North Carolina.

stat.ME

The projected dynamic linear model for time series on the sphere

Time series on the unit n-sphere arise in directional statistics, compositional data analysis, and many scientific fields. There are few models for such data, and the ones that exist suffer from several limitations: they are often computationally challenging to fit, many of them apply only to the circular case of n=2, and they are usually based on families of distributions that are not flexible enough to capture the complexities observed in real data. Furthermore, there is little work on Bayesian methods for spherical time series. To address these shortcomings, we propose a state space model based on the projected normal distribution that can be applied to spherical time series of arbitrary dimension. We describe how to perform fully Bayesian offline inference for this model using a simple and efficient Gibbs sampling algorithm, and we develop a Rao-Blackwellized particle filter to perform online inference for streaming data. In analyses of wind direction and energy market time series, we show that the proposed model outperforms competitors in terms of point, set, and density forecasting.

stat.ME

Bayesian Data Synthesis and the Utility-Risk Trade-Off for Mixed Epidemiological Data

Much of the micro data used for epidemiological studies contain sensitive measurements on real individuals. As a result, such micro data cannot be published out of privacy concerns, rendering any published statistical analyses on them nearly impossible to reproduce. To promote the dissemination of key datasets for analysis without jeopardizing the privacy of individuals, we introduce a cohesive Bayesian framework for the generation of fully synthetic, high dimensional micro datasets of mixed categorical, binary, count, and continuous variables. This process centers around a joint Bayesian model that is simultaneously compatible with all of these data types, enabling the creation of mixed synthetic datasets through posterior predictive sampling. Furthermore, a focal point of epidemiological data analysis is the study of conditional relationships between various exposures and key outcome variables through regression analysis. We design a modified data synthesis strategy to target and preserve these conditional relationships, including both nonlinearities and interactions. The proposed techniques are deployed to create a synthetic version of a confidential dataset containing dozens of health, cognitive, and social measurements on nearly 20,000 North Carolina children.

stat.ME