SearcharxivSearch

arXiv subjects

Gregor Zens

Publications and source records attributed to Gregor Zens.

14 recordsLinked to original sources

Marginal Data Augmentation for Efficient Bayesian Modeling of Counts and Rates with a Demographic Application

Count data models are ubiquitous in many fields, yet Bayesian data augmentation algorithms for such models frequently encounter challenges with Markov chain Monte Carlo efficiency. Posterior simulation is especially demanding when modeling data with a high proportion of zero outcomes. In this paper, we address this issue by introducing a marginal data augmentation approach for semi-parametric Bayesian count data regression models, based on a working parameter that rescales latent outcomes corresponding to zero-count observations. This strategy alleviates the strong posterior dependencies that typically reduce the efficiency of standard data augmentation schemes and leads to substantial gains in sampling efficiency. Synthetic data examples and simulation studies are used to demonstrate the improvements in mixing compared to conventional sampling methods. A demographic application using latent factor analysis to model subnational mortality counts in Austria further underscores the broader applicability of the proposed methodology.

stat.ME

Distributional Forecasting of EU Asylum Applications with Dynamic Multivariate Count Models

We propose a Bayesian framework for joint distributional forecasting of monthly asylum applications across the EU-27. The model decomposes latent application intensities into country-specific random walks and common factors, with idiosyncratic and shared shocks allowed to exhibit heavy tails or stochastic volatility. Using Eurostat data from 2008 to 2026, we evaluate predictive distributions in a rolling out-of-sample exercise, scoring overall distributional accuracy and upper-tail risk. Three findings emerge. First, the preferred specification varies across countries, scoring rules, and horizons, underscoring the need to align models with policy-specific loss functions. Second, joint EU-27 models improve on country-by-country benchmarks, with the largest gains in the upper tail, where preparedness costs are most relevant. Third, random-walk log-intensities provide a useful short-run description of national asylum-application dynamics, especially when combined with flexible innovation dynamics. We conclude by discussing implications for national and EU-level agencies involved in asylum forecasting and preparedness planning.

stat.AP

Probabilistic Estimation of Hidden Migrant Fatalities Along the Central Mediterranean Route

Estimating the number of migrants who die or go missing along dangerous routes such as the Central Mediterranean remains challenging as available records are incomplete. Some incidents are never documented, and fatalities associated with such unobserved incidents are absent from observed totals. We propose a Bayesian approach for probabilistic estimation of total migrant fatalities in such settings. Building on recent developments in multiple-systems estimation, we develop a time-stratified latent-class framework that accommodates missing fatality counts for unobserved incidents. We apply the method to recoded incident-level data from the Missing Migrants Project for the Central Mediterranean route from 2014 to 2025, encompassing 25,712 fatalities across 1,562 incidents. Our model yields 95% credible intervals of 31,922-41,413 fatalities and 2,170-2,952 deadly incidents, indicating that approximately 62%-81% of fatalities and 53%-72% of incidents are reflected in the available data. We estimate that unreported fatalities were concentrated between 2014 and 2016. Furthermore, we document that reporting likelihood increases with incident severity, implying that smaller incidents are most likely to remain undetected. While contingent on modeling assumptions and incomplete data, our method provides a broadly applicable and principled alternative to naive data adjustment methods.

stat.AP

Scalable Variable Selection and Model Averaging for Latent Regression Models Using Approximate Variational Bayes

We propose a fast and theoretically grounded method for Bayesian variable selection and model averaging in latent variable regression models. Our framework addresses three interrelated challenges: (i) intractable marginal likelihoods, (ii) exponentially large model spaces, and (iii) computational costs in large samples. We introduce a novel integrated likelihood approximation based on mean-field variational posterior approximations and establish its asymptotic model selection consistency under broad conditions. To reduce the computational burden, we develop an approximate variational Bayes scheme that fixes the latent regression outcomes for all models at initial estimates obtained under a baseline null model. Despite its simplicity, this approach locally and asymptotically preserves the model-selection behavior of the full variational Bayes approach to first order, at a fraction of the computational cost. Extensive numerical studies - covering probit, tobit, semi-parametric count data models and Poisson log-normal regression - demonstrate accurate inference and large speedups, often reducing runtime from days to hours with comparable accuracy. Applications to real-world data further highlight the practical benefits of the methods for Bayesian inference in large samples and under model uncertainty.

stat.ME

Dynamic Count Models with Flexible Innovation Processes for Irregular Maritime Migration

Motivated by the challenge of analyzing the dynamics of weekly sea border crossings in the Mediterranean (2015-2025) and the English Channel (2018-2025), we develop a Bayesian dynamic framework for modeling heteroskedastic count time series. Building on theoretical considerations and empirical stylized facts, our approach utilizes a Poisson random walk model that allows for heavy-tailed innovations or stochastic volatility dynamics, while incorporating an explicit mechanism to separate structural from sampling zeros. Posterior inference is carried out via a straightforward Markov chain Monte Carlo algorithm. Applying this methodology to Mediterranean and English Channel data, we compare alternative model specifications through a comprehensive out-of-sample forecasting exercise. Using log predictive scores and empirical coverage at predictive quantiles to evaluate each model, we find strong evidence for stochastic volatility in migration innovations. These models deliver the strongest out-of-sample forecasts with empirical coverage close to nominal levels up to the 99th percentile. Our framework can be used to develop risk indicators with direct policy implications for improving governance and preparedness for migration surges. More broadly, the methodology extends to other zero-inflated non-stationary count time series applications, including epidemiological surveillance and public safety incident monitoring.

stat.AP

Low-rank bilinear autoregressive models for three-way criminal activity tensors

Criminal activity data are typically available via a three-way tensor encoding the reported frequencies of different crime categories across time and space. The challenges that arise in the design of interpretable, yet realistic, model-based representations of the complex dependencies within and across these three dimensions have led to an increasing adoption of black-box predictive strategies. While this perspective has proved successful in producing accurate forecasts guiding targeted interventions, the lack of interpretable model-based characterizations of the dependence structures underlying criminal activity tensors prevents from inferring the cascading effects of these interventions across the different dimensions. We address this gap through the design of a low-rank bilinear autoregressive model which achieves comparable predictive performance to black-box strategies, while allowing interpretable inference on the dependence structures of reported criminal activities across crime categories, time and space. This representation incorporates the time dimension via an autoregressive construction that accounts for spatial effects and dependencies among crime categories through a separable low-rank bilinear formulation. When applied to Chicago police reports, the proposed model showcases remarkable predictive performance and also reveals interpretable dependence structures unveiling fundamental crime dynamics. These results facilitate the design of more refined intervention policies informed by the cascading effects of the policy itself.

stat.AP

Bayesian Matrix Factor Models for Demographic Analysis Across Age and Time

Analyzing demographic data collected across multiple populations, time periods, and age groups is challenging due to the interplay of high dimensionality, demographic heterogeneity among groups, and stochastic variability within smaller groups. This paper proposes a Bayesian matrix factor model to address these challenges. By factorizing count data matrices as the product of low-dimensional latent age and time factors, the model achieves a parsimonious representation that mitigates overfitting and remains computationally feasible even when hundreds of populations are involved. Informative priors enforce smoothness in the age factors and allow for the dynamic evolution of the time factors. A straightforward Markov chain Monte Carlo algorithm is developed for posterior inference. Applying the model to Austrian district-level migration data from 2002 to 2023 demonstrates its ability to accurately reconstruct complex demographic processes using only a fraction of the parameters required by conventional demographic factor models. A forecasting exercise shows that the proposed model consistently outperforms standard benchmarks. Beyond statistical demography, the framework holds promise for a wide range of applications involving noisy, heterogeneous, and high-dimensional non-Gaussian matrix-valued data.

stat.AP

Model Uncertainty in Latent Gaussian Models with Univariate Link Function

We consider a class of latent Gaussian models with a univariate link function (ULLGMs). These are based on standard likelihood specifications (such as Poisson, Binomial, Bernoulli, Erlang, etc.) but incorporate a latent normal linear regression framework on a transformation of a key scalar parameter. We allow for model uncertainty regarding the covariates included in the regression. The ULLGM class typically accommodates extra dispersion in the data and has clear advantages for deriving theoretical properties and designing computational procedures. We formally characterize posterior existence under a convenient and popular improper prior and show that ULLGMs inherit the consistency properties from the latent Gaussian model. We propose a simple and general Markov chain Monte Carlo algorithm for Bayesian model averaging in ULLGMs. Simulation results suggest that the framework provides accurate results that are robust to some degree of misspecification. The methodology is successfully applied to measles vaccination coverage data from Ethiopia and to data on bilateral migration flows between OECD countries.

stat.ME

Flexible Bayesian Modelling of Age-Specific Counts in Many Demographic Subpopulations

Analysing age-specific mortality, fertility, and migration patterns is a crucial task in demography with significant policy relevance. In practice, such analysis is challenging when studying a large number of subpopulations, due to small observation counts within groups and increasing demographic heterogeneity between groups. This article proposes a Bayesian model for the joint analysis of age-specific counts in many, potentially small, demographic subpopulations. The model utilizes smooth latent factors to capture common age-specific patterns across subpopulations and facilitates additional information sharing through a hierarchical prior. It provides smoothed estimates of the latent age pattern in each subpopulation, allows testing for heterogeneity, and can be used to assess the impact of covariates on the demographic process. An in-depth case study of age-specific immigration flows to Austria, disaggregated by sex and 155 countries of origin, is discussed. Comparative analysis demonstrates that the model outperforms commonly used benchmark frameworks in both in-sample imputation and out-of-sample predictive exercises.

stat.AP

Ultimate Pólya Gamma Samplers -- Efficient MCMC for possibly imbalanced binary and categorical data

Modeling binary and categorical data is one of the most commonly encountered tasks of applied statisticians and econometricians. While Bayesian methods in this context have been available for decades now, they often require a high level of familiarity with Bayesian statistics or suffer from issues such as low sampling efficiency. To contribute to the accessibility of Bayesian models for binary and categorical data, we introduce novel latent variable representations based on Pólya-Gamma random variables for a range of commonly encountered logistic regression models. From these latent variable representations, new Gibbs sampling algorithms for binary, binomial, and multinomial logit models are derived. All models allow for a conditionally Gaussian likelihood representation, rendering extensions to more complex modeling frameworks such as state space models straightforward. However, sampling efficiency may still be an issue in these data augmentation based estimation frameworks. To counteract this, novel marginal data augmentation strategies are developed and discussed in detail. The merits of our approach are illustrated through extensive simulations and real data applications.

stat.CO

Efficient Bayesian Modeling of Binary and Categorical Data in R: The UPG Package

In this vignette, we introduce the UPG package for efficient Bayesian inference in probit, logit, multinomial logit and binomial logit models. UPG offers a convenient estimation framework for balanced and imbalanced data settings where sampling efficiency is ensured through marginal data augmentation. UPG provides several methods for fast production of output tables and summary plots that are easily accessible to a broad range of users.

stat.CO

A Factor-Augmented Markov Switching (FAMS) Model

This paper investigates the role of high-dimensional information sets in the context of Markov switching models with time varying transition probabilities. Markov switching models are commonly employed in empirical macroeconomic research and policy work. However, the information used to model the switching process is usually limited drastically to ensure stability of the model. Increasing the number of included variables to enlarge the information set might even result in decreasing precision of the model. Moreover, it is often not clear a priori which variables are actually relevant when it comes to informing the switching behavior. Building strongly on recent contributions in the field of factor analysis, we introduce a general type of Markov switching autoregressive models for non-linear time series analysis. Large numbers of time series are allowed to inform the switching process through a factor structure. This factor-augmented Markov switching (FAMS) model overcomes estimation issues that are likely to arise in previous assessments of the modeling framework. More accurate estimates of the switching behavior as well as improved model fit result. The performance of the FAMS model is illustrated in a simulated data example as well as in an US business cycle application.

econ.EM

Bayesian shrinkage in mixture of experts models: Identifying robust determinants of class membership

A method for implicit variable selection in mixture of experts frameworks is proposed. We introduce a prior structure where information is taken from a set of independent covariates. Robust class membership predictors are identified using a normal gamma prior. The resulting model setup is used in a finite mixture of Bernoulli distributions to find homogenous clusters of women in Mozambique based on their information sources on HIV. Fully Bayesian inference is carried out via the implementation of a Gibbs sampler.

econ.EM

Implications of macroeconomic volatility in the Euro area

In this paper we estimate a Bayesian vector autoregressive model with factor stochastic volatility in the error term to assess the effects of an uncertainty shock in the Euro area. This allows us to treat macroeconomic uncertainty as a latent quantity during estimation. Only a limited number of contributions to the literature estimate uncertainty and its macroeconomic consequences jointly, and most are based on single country models. We analyze the special case of a shock restricted to the Euro area, where member states are highly related by construction. We find significant results of a decrease in real activity for all countries over a period of roughly a year following an uncertainty shock. Moreover, equity prices, short-term interest rates and exports tend to decline, while unemployment levels increase. Dynamic responses across countries differ slightly in magnitude and duration, with Ireland, Slovakia and Greece exhibiting different reactions for some macroeconomic fundamentals.

econ.EM