Searcharxiv⌕ Search

arXiv subjects

Johannes Brachem

Publications and source records attributed to Johannes Brachem.

6 recordsLinked to original sources

Reconciling Interpretability with Covariate-Dependent Shape Flexibility in Penalized Transformation Models for Distributional Regression

A central challenge in distributional regression is to allow the shape of the conditional distribution of the response variable to vary flexibly with covariates while retaining directly interpretable effects on its mean and standard deviation. We extend the penalized transformation model (PTM) family into a conditional-shape PTM, which assigns separate structured additive predictors to the conditional mean, standard deviation, and standardized distributional shape beyond location and scale. A covariate-dependent monotone transformation maps the standardized response to a fixed reference distribution, while affine standardization enforces mean zero and variance one for the induced standardized distribution. Thus, the first two predictors remain exactly the conditional mean and standard deviation. The shape predictor accommodates selected linear, nonlinear, group-specific, spatial, and interaction effects; regularization shrinks unsupported departures toward a reference-family location-scale model. We fit the PTM using mini-batched stochastic variational inference with model-aligned Gaussian blocks and staged optimization. In simulations, the PTM recovers smooth mean and standard-deviation effects and a covariate-dependent transition from skewness to bimodality while suppressing unnecessary shape effects. Under a deliberately misspecified design, it remains competitive with a structured additive Dirichlet-process mixture in test-set density and distribution-function accuracy, although all models show undercoverage and both flexible methods miss fine features. Applications to $13{,}425$ Norwegian water-conductivity observations and $1{,}182{,}514$ German daily-temperature observations demonstrate selective group-specific, seasonal, and spatial shape variation. Predictive performance criteria favor the conditional-shape PTM over a fixed-shape PTM and a Gaussian location-scale model.

stat.ME↗

Data-Efficient Generative Modeling of Non-Gaussian Global Climate Fields via Scalable Composite Transformations

Quantifying uncertainty in climate-model output requires characterizing internal variability, often through large ensembles of physical climate-model runs. Since each additional ensemble member is computationally expensive, only limited numbers of replicated fields are typically available under a fixed model configuration and forcing scenario. We propose a data-efficient stochastic generator for the internal variability of global climate fields, specifically designed to overcome these sample-size constraints. The framework targets the distribution of climate-model output under fixed forcing and model physics. Inspired by copula modeling, our approach constructs a highly expressive joint distribution via a composite transformation to a multivariate standard normal space. We combine a nonparametric Bayesian transport map for spatial dependence modeling with flexible, spatially varying marginal models, essential for capturing non-Gaussian behavior and heavy-tailed extremes. These marginals are defined by a parametric model followed by a semi-parametric B-spline correction to capture complex distributional features. The marginal parameters are spatially smoothed using Gaussian-process priors with low-rank approximations, rendering the computational cost linear in the spatial dimension. When applied to global log-precipitation-rate fields under fixed forcing at more than 50,000 grid locations, our stochastic surrogate achieves high fidelity, accurately quantifying the climate distribution's spatial dependence and marginal characteristics, including the tails. Using only 10 training fields, it outperforms a state-of-the-art competitor trained on 80 fields, effectively octupling the computational budget for climate research. We provide a Python implementation at https://github.com/jobrachem/ppptm .

stat.ME↗

Bayesian Penalized Transformation Models: Structured Additive Location-Scale Regression for Arbitrary Conditional Distributions

Penalized transformation models (PTMs) are a semiparametric location-scale regression family that estimate a response's conditional distribution directly from the data, and model the location and scale through structured additive predictors. The core of the model is a monotonically increasing transformation function that relates the response distribution to a reference distribution. The transformation function is equipped with a smoothness prior that regularizes how much the estimated distribution diverges from the reference. PTMs can be seen as a bridge between conditional transformation models and generalized additive models for location, scale and shape. Markov chain Monte Carlo inference for PTMs offers straightforward uncertainty quantification for the conditional distribution as well as for the covariate effects. A simulation study demonstrates the effectiveness of the approach and includes comparisons to many alternative methods. Applications to the Fourth Dutch Growth Study and the Framingham Heart Study illustrate the usage and practical utility. A full-featured implementation is available as a Python library. Supplementary material for this article is available online.

stat.ME↗

Liesel: A Python Framework for Graph-Based Bayesian Modeling and Customizable MCMC with Support for Generalized Additive Models

Liesel is a Python framework for Bayesian model building and posterior computation with dedicated support for generalized additive regression models that is designed to reduce friction in methodological work. The framework consists of three components. The first component, Liesel-Model, represents models as directed acyclic graphs and supports interactive model construction, modification, visualization, prediction, and prior or posterior predictive simulation. The second component, Liesel-Goose, provides a modular MCMC framework based on reusable kernels and supports blocked componentwise sampling as well as user-defined Gibbs and Metropolis--Hastings updates, while leveraging JAX for automatic differentiation, just-in-time compilation, and hardware acceleration. The third component, Liesel-GAM, supplies high-level building blocks for generalized additive regression models and provides functionality for formula-based model specification, summaries, diagnostics, and effect visualization. Together, these parts make Liesel an effective tool for the rapid development, testing, and application of Bayesian models and MCMC algorithms. Liesel's modular architecture allows users to extend the software to models and inference algorithms beyond generalized additive models and MCMC, offering considerable flexibility for a wide spectrum of statistical research.

stat.CO↗

Bayesian structured additive quantile regression for inflated bounded data

Bounded continuous data on the unit interval frequently arise in applied fields and often exhibit a non-negligible proportion of observations at the boundaries. Inflated regression models address this feature by combining a continuous distribution on the unit interval with a discrete component to account for zero- and/or one-inflation. In this paper, we propose a class of Bayesian structured additive quantile regression models for inflated bounded continuous data that accommodates zero- and/or one-inflation. The proposed approach enables direct modeling of both the conditional quantiles of the continuous component and the probabilities of observing zeros and/or ones, with structured additive predictors incorporated in both parts, including nonlinear effects, spatial effects, random effects, and varying-coefficient terms. Posterior inference is carried out using Markov chain Monte Carlo algorithms implemented through the software Liesel, a probabilistic programming framework for semiparametric regression. The practical performance of the proposed models is illustrated through simulation studies and two real-data applications: one analyzing the proportion of traffic-related fatalities across Brazilian municipal districts, and another evaluating speech intelligibility in cochlear implant recipients under different experimental conditions.

stat.ME↗

Graphical Transformation Models

Graphical Transformation Models (GTMs) are introduced as a novel approach to effectively model multivariate data with intricate marginals and complex dependency structures semiparametrically, while maintaining interpretability through the identification of varying conditional independencies. GTMs extend multivariate transformation models by replacing the Gaussian copula with a custom-designed multivariate transformation, offering two major advantages. Firstly, GTMs can capture more complex interdependencies using penalized splines, which also provide an efficient regularization scheme. Secondly, we demonstrate how to approximately regularize GTMs towards pairwise conditional independencies using a lasso penalty, akin to Gaussian graphical models. The model's robustness and effectiveness are validated through simulations, showcasing its ability to accurately learn complex dependencies and identify conditional independencies. Additionally, the model is applied to a benchmark astrophysics dataset, where the GTM demonstrates favorable performance compared to non-parametric vine copulas in learning complex multivariate distributions.

stat.ME↗