SearcharxivSearch

arXiv subjects

David A. Stephens

Publications and source records attributed to David A. Stephens.

At least 19 recordsLinked to original sources

Bootstrap validity in Bayesian semi-parametric models

We discuss Bayesian inference on a low-dimensional targeted parameter in the presence of possibly highly complex nuisance components within the semi-parametric inference framework using an estimating function approach. We obtain a posterior distribution using non-parametric Bayesian methods through the Dirichlet process and the Bayesian bootstrap. We relax the commonly deployed notion of stochastic equicontinuity and develop a framework leading to posterior inference with good frequentist properties, specifically we demonstrate that the posterior distribution is asymptotically Normal and concentrates at the true value of the parameter. We emphasize the specific assumptions that are required to obtain these results, and how relaxing any of them alters the conclusions. We verify the analytical results in simulation.

math.ST

Semi- and non-parametric approaches to individualized treatment regimes in the presence of causal mediation

Individualized treatment rules (ITRs) map an individual patient's characteristics to their recommended treatment value. Typically, the optimal ITR is defined as the rule which maximizes a mean counterfactual outcome; the resulting ITR maximizes the effect of treatment along all causal pathways to the outcome, including indirect pathways through mediating variables. Although maximizing the total effect is often sufficient, explicitly incorporating causal mediation in an ITR analysis has several potential benefits such as enhanced interpretability, and additional flexibility in targeting specific causal pathways. For this purpose, we introduce novel Bayesian semiparametric and nonparametric estimators for conditional mediation effects in the presence of multiple mediators and show how they can be used to estimate optimal ITRs. We demonstrate the proposed methodology via an application to optimal kidney allocation with hepatitis C positive donors.

stat.ME

Not Just How Much, But Where: Decomposing Epistemic Uncertainty into Per-Class Contributions

In safety-critical classification, the cost of failure is often asymmetric, yet Bayesian deep learning summarises epistemic uncertainty with a single scalar, mutual information (MI), that cannot distinguish whether a model's ignorance involves a benign or safety-critical class. We decompose MI into a per-class vector $C_k(x)=\sigma_k^{2}/(2\mu_k)$, with $\mu_k{=}\mathbb{E}[p_k]$ and $\sigma_k^2{=}\mathrm{Var}[p_k]$ across posterior samples. The decomposition follows from a second-order Taylor expansion of the entropy; the $1/\mu_k$ weighting corrects boundary suppression and makes $C_k$ comparable across rare and common classes. By construction $\sum_k C_k \approx \mathrm{MI}$, and a companion skewness diagnostic flags inputs where the approximation degrades. After characterising the axiomatic properties of $C_k$, we validate it on three tasks: (i) selective prediction for diabetic retinopathy, where critical-class $C_k$ reduces selective risk by 34.7\% over MI and 56.2\% over variance baselines; (ii) out-of-distribution detection on clinical and image benchmarks, where $\sum_k C_k$ achieves the highest AUROC and the per-class view exposes asymmetric shifts invisible to MI; and (iii) a controlled label-noise study in which $\sum_k C_k$ shows less sensitivity to injected aleatoric noise than MI under end-to-end Bayesian training, while both metrics degrade under transfer learning. Across all tasks, the quality of the posterior approximation shapes uncertainty at least as strongly as the choice of metric, suggesting that how uncertainty is propagated through the network matters as much as how it is measured.

stat.ML

Semi-parametric Bayesian inference under Neyman orthogonality

The validity of two-step or plug-in inference methods is questioned in the Bayesian framework. We study semi-parametric models where the plug-in of a non-parametrically modelled nuisance component is used. We show that when the nuisance and targeted parameters satisfy a Neyman orthogonal score property, the approach of cutting feedback through a two-step procedure is a valid way of conducting Bayesian inference. Our method relies on a non-parametric Bayesian formulation based on the Dirichlet process and the Bayesian bootstrap. We show that the marginal posterior of the targeted parameter exhibits good frequentist properties despite not accounting for the inferential uncertainty of the nuisance parameter. We adopt this approach in Bayesian causal inference problems where the nuisance propensity score model is estimated to obtain marginal inference for the treatment effect parameter, and demonstrate that a plug-in of the propensity score has a negligible effect on marginal posterior inference for the causal contrast. We investigate the absence of Neyman orthogonality and exploit our findings to show that in conventional two-step procedures, the posterior distribution converges under weaker restrictions than those needed in the frequentist sequel. For a simple family of useful scores, we demonstrate that even in the absence of Neyman orthogonality, the posterior distribution is asymptotically unchanged by the estimation of the nuisance parameter, merely provided the latter estimator is consistent.

math.ST

Posterior Uncertainty for Targeted Parameters in Bayesian Bootstrap Procedures

We propose a general method to carry out a valid Bayesian analysis of a finite-dimensional `targeted' parameter in the presence of a finite-dimensional nuisance parameter. We apply our methods to causal inference based on estimating equations. While much of the literature in Bayesian causal inference has relied on the conventional 'likelihood times prior' framework, a recently proposed method, the 'Linked Bayesian Bootstrap', deviated from this classical setting to obtain valid Bayesian inference using the Dirichlet process and the Bayesian bootstrap. These methods rely on an adjustment based on the propensity score and explain how to handle the uncertainty concerning it when studying the posterior distribution of a treatment effect. We examine theoretically the asymptotic properties of the posterior distribution obtained and show that our proposed method, a generalized version of the 'Linked Bayesian Bootstrap', enjoys desirable frequentist properties. In addition, we show that the credible intervals have asymptotically the correct coverage properties. We discuss the applications of our method to mis-specified and singly-robust models in causal inference.

stat.ME

Singular Bayesian Neural Networks

Bayesian neural networks promise calibrated uncertainty but require $O(mn)$ parameters for standard mean-field Gaussian posteriors. We argue this cost is often unnecessary, particularly when weight matrices exhibit fast singular value decay. By parameterizing weights as $W = AB^{\top}$ with $A \in \mathbb{R}^{m \times r}$, $B \in \mathbb{R}^{n \times r}$, we induce a posterior that is \emph{singular} with respect to the Lebesgue measure, concentrating on the rank-$r$ manifold. This singularity captures structured weight correlations through shared latent factors, geometrically distinct from mean-field's independence assumption. We derive PAC-Bayes generalization bounds whose complexity term scales as $\sqrt{r(m+n)}$ instead of $\sqrt{m n}$, and prove loss bounds that decompose the error into optimization and rank-induced bias using the Eckart-Young-Mirsky theorem. We further adapt recent Gaussian complexity bounds for low-rank deterministic networks to Bayesian predictive means. Empirically, across MLPs, LSTMs, and Transformers on standard benchmarks, our method achieves competitive predictive performance while using up to $33\times$ fewer parameters than 5-member Deep Ensembles. It substantially improves OOD detection and often improves calibration relative to mean-field and perturbation baselines, while Deep Ensembles can still be stronger on in-distribution likelihood-based metrics.

stat.ML

Information Borrowing from Partially Compatible Trajectories for Estimation of Dynamic Treatment Regimes

Dynamic Treatment Regimes (DTRs) provide a systematic framework for optimizing sequential decision-making in chronic disease management, where therapies must adapt to patients' evolving clinical profiles. Inverse probability weighting (IPW) is a cornerstone methodology for estimating regime values from observational data due to its intuitive formulation and established theoretical properties, yet standard IPW estimators face significant limitations, including variance instability and data inefficiency. A fundamental but underexplored source of inefficiency lies in the strict alignment requirement between observed and target treatment trajectories, which fails to account for partial compatibility and discards substantial information from individuals with only minimal deviations from the regime. We propose two novel methodologies that relax the strict inclusion rule through flexible compatibility mechanisms. Both methods provide computationally tractable alternatives that can be easily integrated into existing IPW workflows, offering more efficient approaches to DTR estimation. Theoretical analysis demonstrates that both estimators preserve consistency while achieving superior finite-sample efficiency compared to standard IPW, and comprehensive simulation studies confirm improved stability. We illustrate the practical utility of our methods through an application to HIV treatment data from the AIDS Clinical Trials Group Study 175 (ACTG175).

stat.ME

Individualized treatment regimens under correlated data with multiple outcomes

Precision medicine involves developing individualized treatment regimes (ITRs) which allow for treatment decisions to be tailored to patient characteristics. Naturally, the identification of the optimal regime, that is, the rule which maximizes patient outcomes, is of interest. Several procedures for estimating optimal ITRs from observational data have been proposed; however, relatively few methods exist for estimating optimal ITRs in the presence of competing risks. Previous approaches either target one particular cause of failure, or rely on singly-robust estimators. We propose a novel doubly-robust regression-based method for estimating optimal ITRs which accounts for the uncertainty related to the unobserved cause of failure by averaging over all possible causes, or targeting the most likely cause. Our approach is straightforward to implement, and we demonstrate an extension to incorporate clustering, motivated by the question of for whom kidney transplantation with hepatitis C virus (HCV)-positive donors is safe, using data from the Organ Procurement and Transplantation Network. Our analysis suggests that a large portion of HCV-negative kidney recipients would see their overall survival unchanged if they were instead provided a kidney from an HCV-positive donor. The estimated treatment rules could be used to provide more efficient allocation of HCV-positive kidneys, increasing the donor pool.

stat.ME

Multivariate regression with missing response data for modelling regional DNA methylation QTLs

Identifying genetic regulators of DNA methylation (mQTLs) with multivariate models enhances statistical power, but is challenged by missing data from bisulfite sequencing. Standard imputation-based methods can introduce bias, limiting reliable inference. We propose \texttt{missoNet}, a novel convex estimation framework that jointly estimates regression coefficients and the precision matrix from data with missing responses. By using unbiased surrogate estimators, our three-stage procedure avoids imputation while simultaneously performing variable selection and learning the conditional dependence structure among responses. We establish theoretical error bounds, and our simulations demonstrate that \texttt{missoNet} consistently outperforms existing methods in both prediction and sparsity recovery. In a real-world mQTL analysis of the CARTaGENE cohort, \texttt{missoNet} achieved superior predictive accuracy and false-discovery control on a held-out validation set, identifying known and credible novel genetic associations. The method offers a robust, efficient, and theoretically grounded tool for genomic analyses, and is available as an R package.

stat.ME

Bayesian measurement error modeling of latent time series structure to assess the impact of pollutants on health

The association between levels of air pollution and mortality rate is well-established, but quantifying the magnitude of the effect is sometimes complicated by limitations in the data. In this paper, a joint Bayesian hierarchical model is developed to identify significant predictor factors and quantify their impacts on mortalities. To account for potential measurement error in the pollution data, the observed pollutant levels were treated as noisy proxies for the true exposure, therefore requiring a measurement error structure. We illustrate the developed model by performing an analysis of the association between weekly air pollution levels and cardiovascular and respiratory mortality in Los Angeles (LA) County over five years from January 2018 to December 2022. The purpose of the study was to quantify the impact of the main air pollutants (PM2.5, PM10, SO2, NO2, CO, and O3) and of temperature on weekly mortality from four causes: chronic obstructive pulmonary diseases, pneumonia, heart failures, and malignant neoplasms. The results, supported by Bayesian model comparison criteria, indicated that weekly county-level cause-specific mortalities were significantly associated with certain ranges of air pollutants, and that some of these cause-specific mortalities had significant associations with ambient temperature levels. The analysis revealed that several pollutants appeared to be associated with lower mortality; we interpreted this counterintuitive finding from various perspectives.

stat.AP

Computational Considerations for the Linear Model of Coregionalization

In the last two decades, the linear model of coregionalization (LMC) has been widely used to model multivariate spatial processes. However, it can be a challenging task to conduct likelihood-based inference for such models because of the cubic cost associated with Gaussian likelihood evaluations. Starting from an analogy with matrix normal models, we propose a reformulation of the LMC likelihood that highlights the linear, rather than cubic, computational complexity as a function of the dimension of the response vector. We describe how those simplifications can be exploited in Gaussian hierarchical models. In addition, we propose a new sparsity-inducing approach to the LMC that introduces structural zeros in the coregionalization matrix in an attempt to reduce the number of parameters in a principled and data-driven way. Our reformulation of the LMC likelihood ensures that our sparse approach comes at virtually no additional cost when included in a Markov chain Monte Carlo (MCMC) algorithm. It is shown, on synthetic data, to significantly improve predictive performance. We also apply our methodology to a dataset comprised of air pollutant measurements from the state of California. We investigate the strength of the correlation among the measurements by providing new insights from our sparse method.

stat.ME

Generalized Random Forests using Fixed-Point Trees

We propose a computationally efficient alternative to generalized random forests (GRFs) for estimating heterogeneous effects in large dimensions. While GRFs rely on a gradient-based splitting criterion, which in large dimensions is computationally expensive and unstable, our method introduces a fixed-point approximation that eliminates the need for Jacobian estimation. This gradient-free approach preserves GRF's theoretical guarantees of consistency and asymptotic normality while significantly improving computational efficiency. We demonstrate that our method achieves a speedup of multiple times over standard GRFs without compromising statistical accuracy. Experiments on both simulated and real-world data validate our approach. Our findings suggest that the proposed method is a scalable alternative for localized effect estimation in machine learning and causal inference applications

stat.ML

The impact of directly observed therapy on the efficacy of Tuberculosis treatment: A Bayesian multilevel approach

We propose and discuss a Bayesian procedure to estimate the average treatment effect (ATE) for multilevel observations in the presence of confounding. We focus on situations where the confounders may be latent (e.g., spatial latent effects). This work is motivated by an interest in determining the causal impact of directly observed therapy (DOT) on the successful treatment of Tuberculosis (TB); the available data correspond to individual-level information observed across different cities in a state in Brazil. We focus on propensity score regression and covariate adjustment to balance the treatment (DOT) allocation. We discuss the need to include latent local-level random effects in the propensity score model to reduce bias in the estimation of the ATE. A simulation study suggests that accounting for the multilevel nature of the data with latent structures in both the outcome and propensity score models has the potential to reduce bias in the estimation of causal effects.

stat.ME

Bayesian inference for optimal dynamic treatment regimes in practice

In this work, we examine recently developed methods for Bayesian inference of optimal dynamic treatment regimes (DTRs). DTRs are a set of treatment decision rules aimed at tailoring patient care to patient-specific characteristics, thereby falling within the realm of precision medicine. In this field, researchers seek to tailor therapy with the intention of improving health outcomes; therefore, they are most interested in identifying optimal DTRs. Recent work has developed Bayesian methods for identifying optimal DTRs in a family indexed by $ψ$ via Bayesian dynamic marginal structural models (MSMs) (Rodriguez Duque et al., 2022a); we review the proposed estimation procedure and illustrate its use via the new BayesDTR R package. Although methods in (Rodriguez Duque et al., 2022a) can estimate optimal DTRs well, they may lead to biased estimators when the model for the expected outcome if everyone in a population were to follow a given treatment strategy, known as a value function, is misspecified or when a grid search for the optimum is employed. We describe recent work that uses a Gaussian process ($GP$) prior on the value function as a means to robustly identify optimal DTRs (Rodriguez Duque et al., 2022b). We demonstrate how a $GP$ approach may be implemented with the BayesDTR package and contrast it with other value-search approaches to identifying optimal DTRs. We use data from an HIV therapeutic trial in order to illustrate a standard analysis with these methods, using both the original observed trial data and an additional simulated component to showcase a longitudinal (two-stage DTR) analysis.

stat.ME

A Bayesian Non-Stationary Heteroskedastic Time Series Model for Multivariate Critical Care Data

We propose a multivariate GARCH model for non-stationary health time series by modifying the variance of the observations of the standard state space model. The proposed model provides an intuitive way of dealing with heteroskedastic data using the conditional nature of state space models. We follow the Bayesian paradigm to perform the inference procedure. In particular, we use Markov chain Monte Carlo methods to obtain samples from the resultant posterior distribution. Due to the natural temporal correlation structure induced on model parameters, we use the forward filtering backward sampling algorithm to efficiently obtain samples from the posterior distribution. The proposed model also handles missing data in a fully Bayesian fashion. We validate our model on synthetic data, and then use it to analyze a data set obtained from an intensive care unit in a Montreal hospital. We further show that our proposed models offer better performance, in terms of WAIC, than standard state space models. The proposed model provides a new way to model multivariate heteroskedastic non-stationary time series data and the simplicity in applying the WAIC allows us to compare competing models.

stat.ME

A time-dependent Poisson-Gamma model for recruitment forecasting in multicenter studies

Forecasting recruitments is a key component of the monitoring phase of multicenter studies. One of the most popular techniques in this field is the Poisson-Gamma recruitment model, a Bayesian technique built on a doubly stochastic Poisson process. This approach is based on the modeling of enrollments as a Poisson process where the recruitment rates are assumed to be constant over time and to follow a common Gamma prior distribution. However, the constant-rate assumption is a restrictive limitation that is rarely appropriate for applications in real studies. In this paper, we illustrate a flexible generalization of this methodology which allows the enrollment rates to vary over time by modeling them through B-splines. We show the suitability of this approach for a wide range of recruitment behaviors in a simulation study and by estimating the recruitment progression of the Canadian Co-infection Cohort (CCC).

stat.ME

The role of exchangeability in causal inference

Though the notion of exchangeability has been discussed in the causal inference literature under various guises, it has rarely taken its original meaning as a symmetry property of probability distributions. As this property is a standard component of Bayesian inference, we argue that in Bayesian causal inference it is natural to link the causal model, including the notion of confounding and definition of causal contrasts of interest, to the concept of exchangeability. Here we propose a probabilistic between-group exchangeability property as an identifying condition for causal effects, relate it to alternative conditions for unconfounded inferences (commonly stated using potential outcomes) and define causal contrasts in the presence of exchangeability in terms of posterior predictive expectations for further exchangeable units. While our main focus is on a point treatment setting, we also investigate how this reasoning carries over to longitudinal settings.

stat.ME

Targeting functional parameters with semiparametric Bayesian inference

Typical Bayesian inference requires parameter identification via likelihood parameterization, which has invited criticism for being less flexible than the Frequentist framework and subject to misspecification. Though misspecification may be avoided by functional parameter inference under a nonparametric model space, there does not exist a flexible Bayesian semiparametric model that would allow full control over the marginal prior over any general functional parameter. We present the technique of $θ$-augmentation which helps us manipulate nonparametric models into semiparametric ones that directly target any functional parameter. The method allows Bayesian probabilistic statements to be drawn for any estimator that is defined as a functional of the empirical distribution without requiring a likelihood function, thus providing a path to Bayesian analysis in problems like causal inference and censoring where there do not exist well-accepted likelihood functions.

stat.ME