Searcharxiv⌕ Search

arXiv subjects

Adrian E. Raftery

Publications and source records attributed to Adrian E. Raftery.

At least 19 recordsLinked to original sources

Estimating Climate Sensitivity Using Bayesian Model Averaging for CMIP Models

The Transient Climate Response to cumulative CO2 Emissions (TCRE) is a key metric for linking greenhouse gas emissions to global temperature change and informing climate policy. However, extant estimates of the TCRE often depend on subjective model selection or assumed sensitivity ranges, with limited validation against observed data. We develop a fully statistical data-driven approach using a Bayesian Model Averaging (BMA) approach to estimate the TCRE. This uses 37 climate models from the Coupled Model Intercomparison Project phase 6 (CMIP6), weighted according to their consistency with observed temperature data. Compared to the Intergovernmental Panel on Climate Change (IPCC)'s Sixth Assessment Report (AR6) TCRE estimate, our BMA approach yields a very likely range (90% interval) that overlaps substantially with that from the AR6, but with a higher mean and a lower standard deviation. The resulting projections of global temperature change to 2100 show somewhat higher warming and less uncertainty than some current methods. The use of statistical modeling methods makes it easier to validate the approach and to partition the uncertainty according to its sources. Out-of-sample predictive validation shows the method to be well calibrated. Variance decomposition shows model parameter uncertainty to be a main source of projection variance.

stat.ME↗

Simulation-consistent Estimation of the Marginal Likelihood for Block Models

We propose a methodology for computing marginal likelihoods for block models. The proposed estimator computes the marginal likelihood from Markov chain Monte Carlo (MCMC) samples and is simulation-consistent, even when the size of the dataset is fixed. Moreover, it is asymptotically normal, of finite variance, invariant to label switching and can be computed efficiently, even for models with an arbitrarily large number of components. We evaluate the method through simulation studies in settings where the true marginal likelihood is available analytically. Finally, we apply the approach to a social network dataset based on the 2023 United Nations Climate Change Conference (COP28) and discuss the resulting insights.

stat.ME↗

Comparing Variable Selection and Model Averaging Methods for Logistic Regression

Model uncertainty is a central challenge in statistical models for binary outcomes such as logistic regression, arising when it is unclear which predictors should be included in the model. Many methods have been proposed to address this issue for logistic regression, but their relative performance under realistic conditions remains poorly understood. We therefore conducted a preregistered, simulation-based comparison of 28 established methods for variable selection and inference under model uncertainty, using 11 empirical datasets spanning a range of sample sizes and number of predictors, in cases both with and without separation. We found that Bayesian model averaging (BMA) methods based on g-priors, particularly g = max(n, p^2), show the strongest overall performance when separation is absent. When separation occurs, penalized likelihood approaches, especially the LASSO, provide the most stable results, while BMA with the local empirical Bayes (EB-local) prior is competitive in both situations. These findings offer practical guidance for applied researchers on how to effectively address model uncertainty in logistic regression in modern empirical and machine learning research.

stat.ME↗

Reversible Jump MCMC With No Regrets: Bayesian Variable Selection Using Mixtures of Mutually Singular Distributions

Bayesian variable selection requires sampling from a posterior distribution that combines discrete model indicators with continuously varying parameters, a challenge often addressed through reversible jump Markov chain Monte Carlo (RJMCMC). Despite its generality, RJMCMC is widely regarded as difficult to design and implement correctly. We present mixtures of mutually singular (MoMS) distributions as a transparent alternative in which competing models are represented within a single fixed-dimensional parameter space partitioned into mutually singular subspaces. We show that this formulation reproduces the exact spike-and-slab interpretation of Bayesian variable selection and that, under appropriate constructions, MoMS and RJMCMC share the same Metropolis--Hastings acceptance probability. On a benchmark dataset with ten predictors, both methods recover posterior inclusion probabilities that match full enumeration, while MoMS achieves comparable or superior effective sample size per second relative to a carefully engineered RJMCMC scheme. We further illustrate the approach in a mixed-effects logistic regression for a sleep-and-memory experiment and in factor-loading selection for a multidimensional generalized partial credit model. Together, these results show that Bayesian variable selection can be carried out within standard fixed-dimensional Markov chain Monte Carlo methodology -- without regret.

stat.ME↗

Bayesian Projection of Extant Refugee and Asylum Seeker Populations

Estimates of future migration patterns are of broad interest in demography. Forced migration, including refugee and asylum seekers, plays an important role in overall migration patterns, but is notoriously difficult to forecast. Focusing on refugees and asylum seekers, we propose a modeling pipeline based on Bayesian hierarchical time-series modeling for projecting refugee population official statistics by country of origin using data from the United Nations High Commissioner for Refugees (UNHCR). Our approach is based on a conceptual model of refugee and asylum seeker populations following growth and decline phases, separated by a peak. The growth and decline phases are modeled by logistic growth and decline through an interrupted logistic process model. We evaluate our method through a set of validation exercises that show it has good performance for forecasts at 1, 5, and 10 year horizons, and we present projections for 35 countries of origin of large refugee and asylum seeker population.

stat.AP↗

Easily Computed Marginal Likelihoods for Multivariate Mixture Models Using the THAMES Estimator

We present a new version of the truncated harmonic mean estimator (THAMES) for univariate or multivariate mixture models. The estimator computes the marginal likelihood from Markov chain Monte Carlo (MCMC) samples, is consistent, asymptotically normal and of finite variance. In addition, it is invariant to label switching, does not require posterior samples from hidden allocation vectors, and is easily approximated, even for an arbitrarily high number of components. Its computational efficiency is based on an asymptotically optimal ordering of the parameter space, which can in turn be used to provide useful visualisations. We test it in simulation settings where the true marginal likelihood is available analytically. It performs well against state-of-the-art competitors, even in multivariate settings with a high number of components. We demonstrate its utility for inference and model selection on univariate and multivariate data sets.

stat.ME↗

Multiple Imputation of Hierarchical Nonlinear Time Series Data with an Application to School Enrollment Data

International comparisons of hierarchical time series data sets based on survey data, such as annual country-level estimates of school enrollment rates, can suffer from large amounts of missing data due to differing coverage of surveys across countries and across times. A popular approach to handling missing data in these settings is through multiple imputation, which can be especially effective when there is an auxiliary variable that is strongly predictive of and has a smaller amount of missing data than the variable of interest. However, standard methods for multiple imputation of hierarchical time series data can perform poorly when the auxiliary variable and the variable of interest have a nonlinear relationship. Performance can also suffer if the multiple imputations are used to estimate an analysis model that makes different assumptions about the data compared to the imputation model, leading to uncongeniality between analysis and imputation models. We propose a Bayesian method for multiple imputation of hierarchical nonlinear time series data that uses a sequential decomposition of the joint distribution and incorporates smoothing splines to account for nonlinear relationships between variables. We compare the proposed method with existing multiple imputation methods through a simulation study and an application to secondary school enrollment data. We find that the proposed method can lead to substantial performance increases for estimation of parameters in uncongenial analysis models and for prediction of individual missing values.

stat.ME↗

US COVID-19 school closure was not cost-effective, but other measures were

Non-pharmaceutical interventions (NPIs) in response to the COVID-19 pandemic necessitated a trade-off between the health impacts of viral spread and the social and economic costs of restrictions. We conduct a cost-effectiveness analysis of NPI policies enacted at the state level in the United States in 2020. Although school closures reduced viral transmission, their social impact in terms of student learning loss was too costly, depriving the nation of \$2 trillion (USD2020), conservatively, in future GDP. Moreover, this marginal trade-off between school closure and COVID deaths was not inescapable: a combination of other measures would have been enough to maintain similar or lower mortality rates without incurring such profound learning loss. Optimal policies involve consistent implementation of mask mandates, public test availability, contact tracing, social distancing orders, and reactive workplace closures, with no closure of schools beyond the usual 16 weeks of break per year. Their use would have reduced the gross impact of the pandemic in the US in 2020 from \$4.6 trillion to \$1.9 trillion and, with high probability, saved over 100,000 lives. Our results also highlight the need to address the substantial global learning deficit incurred during the pandemic.

stat.AP↗

Forecasting Net Migration By Age: The Flow-Difference Approach

Most population projection models require age-specific information on net migration totals as a key demographic component of population change. Existing methods for predicting future patterns of net migration by age have proven inadequate. The main reason is that methods applied to model net migration are unable to distinguish factors influencing the inflows from those influencing the outflows. In this paper, we develop two flow-difference methods to produce age-specific forecasts of net migration for counties in the Washington State. One uses a deterministic approach; the other uses a Bayesian approach and includes measures of uncertainty. Both methods model the age-specific flows of in-migration and out-migration to derive age-specific net migration. By including models for in-migration and out-migration, even in the absence of data on such flows, the resulting net migration predictions are greatly improved over existing methods that only model the net migration totals. The estimation intervals from the Bayesian flow-difference method are found to be well calibrated, while the other approaches do not yield such intervals. The implications for future county-level population projections in Washington State are shown.

stat.AP↗

A Structured Estimator for large Covariance Matrices in the Presence of Pairwise and Spatial Covariates

We consider the problem of estimating a high-dimensional covariance matrix from a small number of observations when covariates on pairs of variables are available and the variables can have spatial structure. This is motivated by the problem arising in demography of estimating the covariance matrix of the total fertility rate (TFR) of 195 different countries when only 11 observations are available. We construct an estimator for high-dimensional covariance matrices by exploiting information about pairwise covariates, such as whether pairs of variables belong to the same cluster, or spatial structure of the variables, and interactions between the covariates. We reformulate the problem in terms of a mixed effects model. This requires the estimation of only a small number of parameters, which are easy to interpret and which can be selected using standard procedures. The estimator is consistent under general conditions, and asymptotically normal. It works if the mean and variance structure of the data is already specified or if some of the data are missing. We assess its performance under our model assumptions, as well as under model misspecification, using simulations. We find that it outperforms several popular alternatives. We apply it to the TFR dataset and draw some conclusions.

stat.ME↗

Bringing Age Back In: Accounting for Population Age Distribution in Forecasting Migration

The link between age and migration propensity is long established, but existing models of country-level net migration ignore the effect of population age distribution on past and projected migration rates. We propose a method to estimate and forecast international net migration rates for the 200 most populous countries, taking account of changes in population age structure. We use age-standardized estimates of country-level net migration rates and in-migration rates over quinquennial periods from 1990 through 2020 to decompose past net migration rates into in-migration rates and out-migration rates. We then recalculate historic migration rates on a scale that removes the influence of the population age distribution. This is done by scaling past and projected migration rates in terms of a reference population and period. We show that this can be done very simply, using a quantity we call the migration age structure index (MASI). We use a Bayesian hierarchical model to generate joint probabilistic forecasts of total and age- and sex- specific net migration rates over five-year periods for all countries from 2020 through 2100. We find that accounting for population age structure in historic and forecast net migration rates leads to narrower prediction intervals by the end of the century for most countries. Also, applying a Rogers & Castro-like migration age schedule to migration outflows reduces uncertainty in population pyramid forecasts. Finally, accounting for population age structure leads to less out-migration among countries with rapidly aging populations that are forecast to contract most rapidly by the end of the century. This leads to less drastic population declines than are forecast without accounting for population age structure.

stat.AP↗

Interview with Adrian Raftery

Professor Adrian E. Raftery is the Boeing International Professor of Statistics and Sociology, and an adjunct professor of Atmospheric Sciences, at the University of Washington in Seattle. He was born in Dublin, Ireland, and obtained a B.A. in Mathematics and an M.Sc. in Statistics and Operations Research at Trinity College Dublin. He obtained a doctorate in mathematical statistics from the Université Pierre et Marie Curie under the supervision of Paul Deheuvels. He was a lecturer in statistics at Trinity College Dublin, and then an associate and full professor of statistics and sociology at the University of Washington. He was the founding Director of the Center for Statistics and Social Sciences. Professor Raftery has published over 200 articles in peer-reviewed statistical, sociological and other journals. His research focuses on Bayesian model selection and Bayesian model averaging, model-based clustering, inference for deterministic models, and the development of new statistical methods for demography, sociology, and the environmental and health sciences. He is a member of the United States National Academy of Sciences, a Fellow of the American Academy of Arts and Sciences, an Honorary Member of the Royal Irish Academy, a member of the Washington State Academy of Sciences, a Fellow of the American Statistical Association, a Fellow of the Institute of Mathematical Statistics, and an elected Member of the Sociological Research Association. He has won multiple awards for his research. He was Coordinating and Applications Editor of the Journal of the American Statistical Association and Editor of Sociological Methodology. He was identified as the world's most cited researcher in mathematics for the period 1995-2005. Thirty-three students have obtained Ph.D.'s working under Raftery's supervision, of whom 21 hold or have held tenure-track university faculty positions.

stat.OT↗

Bayesian Hyperbolic Multidimensional Scaling

Multidimensional scaling (MDS) is a widely used approach to representing high-dimensional, dependent data. MDS works by assigning each observation a location on a low-dimensional geometric manifold, with distance on the manifold representing similarity. We propose a Bayesian approach to multidimensional scaling when the low-dimensional manifold is hyperbolic. Using hyperbolic space facilitates representing tree-like structures common in many settings (e.g. text or genetic data with hierarchical structure). A Bayesian approach provides regularization that minimizes the impact of measurement error in the observed data and assesses uncertainty. We also propose a case-control likelihood approximation that allows for efficient sampling from the posterior distribution in larger data settings, reducing computational complexity from approximately $O(n^2)$ to $O(n)$. We evaluate the proposed method against state-of-the-art alternatives using simulations, canonical reference datasets, Indian village network data, and human gene expression data.

stat.ME↗

Easily Computed Marginal Likelihoods from Posterior Simulation Using the THAMES Estimator

We propose an easily computed estimator of marginal likelihoods from posterior simulation output, via reciprocal importance sampling, combining earlier proposals of DiCiccio et al (1997) and Robert and Wraith (2009). This involves only the unnormalized posterior densities from the sampled parameter values, and does not involve additional simulations beyond the main posterior simulation, or additional complicated calculations. It is unbiased for the reciprocal of the marginal likelihood, consistent, has finite variance, and is asymptotically normal. It involves one user-specified control parameter, and we derive an optimal way of specifying this. We illustrate it with several numerical examples.

stat.ME↗

Latent Position Network Models

In this chapter, we present a review of latent position models for networks. We review the recent literature in this area and illustrate the basic aspects and properties of this modeling framework. Through several illustrative examples we highlight how the latent position model is able to capture important features of observed networks. We emphasize how the canonical design of this model has made it popular thanks to its ability to provide interpretable visualizations of complex network interactions. We outline the main extensions that have been introduced to this model, illustrating its flexibility and applicability.

stat.ME↗

Probabilistic Estimation and Projection of the Annual Total Fertility Rate Accounting for Past Uncertainty: A Major Update of the bayesTFR R Package

The bayesTFR package for R provides a set of functions to produce probabilistic projections of the total fertility rates (TFR) for all countries, and is widely used, including as part of the basis for the UN's official population projections for all countries. Liu and Raftery (2020) extended the theoretical model by adding a layer that accounts for the past TFR estimation uncertainty. A major update of bayesTFR implements the new extension. Moreover, a new feature of producing annual TFR estimation and projections extends the existing functionality of estimating and projecting for five-year time periods. An additional autoregressive component has been developed in order to account for the larger autocorrelation in the annual version of the model. This article summarizes the updated model, describes the basic steps to generate probabilistic estimation and projections under different settings, compares performance, and provides instructions on how to summarize, visualize and diagnose the model results.

stat.AP↗

Estimating SARS-CoV-2 Infections from Deaths, Confirmed Cases, Tests, and Random Surveys

There are many sources of data giving information about the number of SARS-CoV-2 infections in the population, but all have major drawbacks, including biases and delayed reporting. For example, the number of confirmed cases largely underestimates the number of infections, deaths lag infections substantially, while test positivity rates tend to greatly overestimate prevalence. Representative random prevalence surveys, the only putatively unbiased source, are sparse in time and space, and the results come with a big delay. Reliable estimates of population prevalence are necessary for understanding the spread of the virus and the effects of mitigation strategies. We develop a simple Bayesian framework to estimate viral prevalence by combining the main available data sources. It is based on a discrete-time SIR model with time-varying reproductive parameter. Our model includes likelihood components that incorporate data of deaths due to the virus, confirmed cases, and the number of tests administered on each day. We anchor our inference with data from random sample testing surveys in Indiana and Ohio. We use the results from these two states to calibrate the model on positive test counts and proceed to estimate the infection fatality rate and the number of new infections on each day in each state in the USA. We estimate the extent to which reported COVID cases have underestimated true infection counts, which was large, especially in the first months of the pandemic. We explore the implications of our results for progress towards herd immunity.

stat.AP↗

The vote Package: Single Transferable Vote and Other Electoral Systems in R

We describe the vote package in R, which implements the plurality (or first-past-the-post), two-round runoff, score, approval and single transferable vote (STV) electoral systems, as well as methods for selecting the Condorcet winner and loser. We emphasize the STV system, which we have found to work well in practice for multi-winner elections with small electorates, such as committee and council elections, and the selection of multiple job candidates. For single-winner elections, the STV is also called instant runoff voting (IRV), ranked choice voting (RCV), or the alternative vote (AV) system. The package also implements the STV system with equal preferences, for the first time in a software package, to our knowledge. It also implements a new variant of STV, in which a minimum number of candidates from a specified group are required to be elected. We illustrate the package with several real examples.

stat.CO↗