SearcharxivSearch

arXiv subjects

Christel Faes

Publications and source records attributed to Christel Faes.

15 recordsLinked to original sources

Integrating Temporal Disaggregation and Distributed Lag Nonlinear Models for Bayesian Spatio-Temporal Disease Mapping with High-Resolution Environmental Exposures

Environmental conditions are major drivers of malaria transmission, but epidemiological analyses are often constrained by temporal misalignment between health outcomes reported at coarse time scales and environmental exposures available at finer resolutions. Conventional approaches aggregate environmental data to match health outcomes, potentially obscuring delayed and nonlinear relationships. We propose a Bayesian spatio-temporal framework that addresses this limitation through a latent daily disease process linked to observed monthly malaria counts by temporal disaggregation. The framework integrates distributed lag nonlinear models for climatic effects, spatio-temporal random effects, and intervention covariates within a unified hierarchical model. The methodology was applied to malaria surveillance data from 161 districts in Mozambique between 2017 and 2024, integrating temperature, precipitation, relative humidity, vegetation, elevation, and malaria interventions. Compared with a conventional monthly model, the proposed framework improved predictive accuracy and uncertainty quantification while exploiting the temporal resolution of environmental data. Estimated relationships showed nonlinear associations between climatic variability and malaria incidence, including an optimal temperature range, increasing risk with positive vegetation anomalies, and nonlinear precipitation effects. By avoiding temporal aggregation of environmental exposures, the framework provides a flexible approach for investigating delayed environmental effects from routine surveillance data and can be extended to other environmentally sensitive diseases with mismatched temporal resolutions.

stat.AP

Laplacian-P-splines for shared Gamma frailty models applied to clustered right-censored time-to-event data

Shared frailty models have been proposed to accommodate unmeasured cluster-specific risk factors through the inclusion of a common latent frailty term. Among possible frailty distributions, the Gamma distribution is appealing due to its non-negativity, flexibility, and algebraic tractability leading to closed-form marginal survival or hazard function expressions. Under the Bayesian paradigm, the posterior distributions of model parameters are usually explored with computationally intensive procedures relying on Markov chain Monte Carlo sampling. As an alternative, Laplacian-P-splines (LPS) provide a flexible and sampling-free alternative by relying on Gaussian approximations of the posterior target distributions. In this model class, analytical formulas are obtained for the gradient and Hessian, yielding a computationally efficient inference scheme for estimation of model parameters with a natural way of quantifying uncertainty. This article extends the LPS toolbox to the inclusion of shared Gamma frailty models for clustered time-to-event data. We assess the finite-sample performance of the LPS estimation procedure through an extensive simulation study and compare estimates with those obtained using penalized partial likelihood estimation, without specification of the baseline hazard, and with the variance of the frailty term being estimated using profile likelihood. Finally, the proposed LPS estimation method is exemplified using three publicly available biomedical datasets on: (i) recurrent infections in children, (ii) cancer prevention, and (iii) kidney transplantation.

stat.ME

Federated generalized linear mixed models based on one-time shared summary statistics

Data privacy has increasingly become a daunting challenge because it limits data availability, which is essential in estimating statistical models such as generalized linear mixed models. Access to personal data often involves considerable time, effort, and paperwork, which can impede research progress and collaboration. Existing approaches that do not use individual-level data for model estimation are either prone to ecological bias, cannot handle heterogeneity, or require iterative communication. In this paper, we propose an approach to estimate generalized linear mixed models based on summary statistics shared only once. We used linear, logistic, and Poisson mixed models as examples to demonstrate the methodology. Our strategy involves generating pseudo-data whose summary statistics match those of the actual but unavailable data. These pseudo-data are then used for model estimation instead of the actual data. The estimates we achieve are identical (up to the third decimal place) to those derived from actual data and have similar bias, coverage, and prediction performance. Communication and resource efficiency distinguish our approach from existing methods.

stat.ME

Spatially varying distributed lag non-linear models using Laplacian P-splines

Although distributed lag non-linear models (DLNMs) are commonly used to quantify delayed and non-linear exposure-response relationships, most existing applications assume that these relationships are constant across space. However, in many geographical and environmental studies, local characteristics vary substantially across areas, making a spatially varying effect more realistic. Extending DLNMs to allow for spatial heterogeneity remains challenging, and only a limited number of modelling strategies have been proposed in literature. The most popular extension is a two-stage meta-analysis approach, which requires sufficiently large sample sizes at each location. Therefore, its usefulness is limited when working with sparse count data in small area data analyses. Although a number of alternative one-stage approaches have been introduced, their computational burden restricts their applicability in real-life data applications. In this paper, we introduce a computationally efficient Bayesian one-stage spatially-varying DLNM for count data. We define four model variants, differing in the assumed spatial dependence structure and the flexibility of the DLNM spline specification. To address the computational burden typically associated with these flexible models, we use Laplace approximations, offering an efficient alternative to classically used Markov Chain Monte Carlo (MCMC) approaches. Model comparison criteria are provided to facilitate the selection of a suitable model in a real-life data application. The proposed methods are evaluated through simulation studies, and their practical usefulness is illustrated through a real-life data application, investigating the temperature-mortality relationship in every municipality of Sicily, Italy.

stat.ME

Distributed lag non-linear models with spatial effect modification using Laplacian P-splines

Distributed lag non-linear models (DLNMs) are a popular approach to flexibly model the effect of time-delayed exposures. Classical DLNMs specify a common exposure-lag-response relationship across geographical areas. However, this relationship might be altered by an effect modifier that differs between spatial units. Although some methods have been proposed to account for effect modification, their applicability is context-dependent. For example, a meta-analysis can account for heterogeneity between groups, but this technique requires sufficiently large study groups. This limitation is particularly relevant when working with count data, where small numbers of events are often encountered. In this paper, we review existing methods that allow for spatial effect modification for count-based outcomes and propose a Bayesian DLNM alternative method that accounts for the modifier through flexible interaction effects. Through the use of Laplacian P-splines, we provide a computationally fast estimation procedure by avoiding the use of classical Markov Chain Monte Carlo (MCMC) approaches. The performance of the different methods is evaluated through simulation studies. Moreover, the practical applicability of our proposed method is showcased through a data application, containing daily temperature and mortality count data in 87 Italian cities.

stat.ME

Backcasting biodiversity at high spatiotemporal resolution using flexible site-occupancy models for opportunistically sampled citizen science data

For many taxonomic groups, online biodiversity portals used by naturalists and citizen scientists constitute the primary source of distributional information. Over the last decade, site-occupancy models have been advanced as a promising framework to analyse such loosely structured, opportunistically collected datasets. Current approaches often ignore important aspects of the detection process and do not fully capitalise on the information present in these datasets, leaving opportunities for fine-grained spatiotemporal backcasting untouched. We propose a flexible Bayesian spatiotemporal site-occupancy model that aims to mimic the data-generating process that underlies common citizen science datasets sourced from public biodiversity portals, and yields rich biological output. We illustrate the use of the model to a dataset containing over 3M butterfly records in Belgium, collected through the citizen science data portal Observations.be. We show that the proposed approach enables retrospective predictions on the occupancy of species through time and space at high resolution, as well as inference on inter-annual distributional trends, range dynamics, habitat preferences, phenological patterns, detection patterns and observer heterogeneity. The proposed model can be used to increase the value of opportunistically collected data by naturalists and citizen scientists, and can aid the understanding of spatiotemporal dynamics of species for which rigorously collected data are absent or too costly to collect.

stat.AP

A Time-Series Model for Areal Data Using Area-Specific Gaussian Processes with Spatially Correlated Hyperparameters

In many applied settings, areal data are observed repeatedly over long time periods, as commonly occurs in infectious disease surveillance and environmental or demographic monitoring. Accurate characterization of local temporal dynamics and uncertainty is important for monitoring disease trends, identifying local changes, and supporting public health decision-making. Traditional spatio-temporal models generally represent spatial, temporal, and space-time interaction components through structured random effects acting on the latent outcome process. We propose a Bayesian spatio-temporal hierarchical framework in which temporal dynamics are modeled using area-specific Gaussian processes, while spatial dependence is introduced through spatially correlated Gaussian-process covariance hyperparameters. This allows neighboring regions to share information about the characteristics of their temporal dependence while retaining area-specific temporal trajectories, providing an alternative representation of spatio-temporal dependence. Inference is performed using Markov chain Monte Carlo methods. The approach is illustrated using monthly malaria incidence data from three Mozambican provinces and evaluated using the Root Mean Squared Error, Continuous Ranked Probability Score, empirical coverage probability, and credible interval width. Compared with established spatio-temporal models, the proposed framework achieves competitive predictive accuracy and better-calibrated predictive uncertainty. Spatially structuring the temporal hyperparameters also improves predictive interval calibration relative to an equivalent model with independent temporal processes, demonstrating the value of borrowing spatial information on temporal dependence and supporting this formulation as a competitive alternative to conventional outcome-level spatial smoothing for spatio-temporal areal data in disease surveillance.

stat.ME

A Bayesian Geoadditive Model for Spatial Disaggregation

We present a novel Bayesian spatial disaggregation model for count data, providing fast and flexible inference at high resolution. First, it incorporates non-linear covariate effects using penalized splines, a flexible approach that is not typically included in existing spatial disaggregation methods. Additionally, it employs a spline-based low-rank kriging approximation for modeling spatial dependencies. The use of Laplace approximation provides computational advantages over traditional Markov Chain Monte Carlo (MCMC) approaches, facilitating scalability to large datasets. We explore two estimation strategies: one using the exact likelihood and another leveraging a spatially discrete approximation for enhanced computational efficiency. Simulation studies demonstrate that both methods perform well, with the approximate method offering significant computational gains. We illustrate the applicability of our model by disaggregating disease rates in the United Kingdom and Belgium, showcasing its potential for generating high-resolution risk maps. By combining flexibility in covariate modeling, computational efficiency and ease of implementation, our approach offers a practical and effective framework for spatial disaggregation.

stat.ME

Distributed lag non-linear models with Laplacian-P-splines for analysis of spatially structured time series

Distributed lag non-linear models (DLNM) have gained popularity for modeling nonlinear lagged relationships between exposures and outcomes. When applied to spatially referenced data, these models must account for spatial dependence, a challenge that has yet to be thoroughly explored within the penalized DLNM framework. This gap is mainly due to the complex model structure and high computational demands, particularly when dealing with large spatio-temporal datasets. To address this, we propose a novel Bayesian DLNM-Laplacian-P-splines (DLNM-LPS) approach that incorporates spatial dependence using conditional autoregressive (CAR) priors, a method commonly applied in disease mapping. Our approach offers a flexible framework for capturing nonlinear associations while accounting for spatial dependence. It uses the Laplace approximation to approximate the conditional posterior distribution of the regression parameters, eliminating the need for Markov chain Monte Carlo (MCMC) sampling, often used in Bayesian inference, thus improving computational efficiency. The methodology is evaluated through simulation studies and applied to analyze the relationship between temperature and mortality in London.

stat.ME

Federated mixed effects logistic regression based on one-time shared summary statistics

Upholding data privacy especially in medical research has become tantamount to facing difficulties in accessing individual-level patient data. Estimating mixed effects binary logistic regression models involving data from multiple data providers like hospitals thus becomes more challenging. Federated learning has emerged as an option to preserve the privacy of individual observations while still estimating a global model that can be interpreted on the individual level, but it usually involves iterative communication between the data providers and the data analyst. In this paper, we present a strategy to estimate a mixed effects binary logistic regression model that requires data providers to share summary statistics only once. It involves generating pseudo-data whose summary statistics match those of the actual data and using these into the model estimation process instead of the actual unavailable data. Our strategy is able to include multiple predictors which can be a combination of continuous and categorical variables. Through simulation, we show that our approach estimates the true model at least as good as the one which requires the pooled individual observations. An illustrative example using real data is provided. Unlike typical federated learning algorithms, our approach eliminates infrastructure requirements and security issues while being communication efficient and while accounting for heterogeneity.

stat.ME

Linear mixed modelling of federated data when only the mean, covariance, and sample size are available

In medical research, individual-level patient data provide invaluable information, but the patients' right to confidentiality remains of utmost priority. This poses a huge challenge when estimating statistical models such as linear mixed models, which is an extension of linear regression models that can account for potential heterogeneity whenever data come from different data providers. Federated learning algorithms tackle this hurdle by estimating parameters without retrieving individual-level data. Instead, iterative communication of parameter estimate updates between the data providers and analyst is required. In this paper, we propose an alternative framework to federated learning algorithms for fitting linear mixed models. Specifically, our approach only requires the mean, covariance, and sample size of multiple covariates from different data providers once. Using the principle of statistical sufficiency within the framework of likelihood as theoretical support, this proposed framework achieves estimates identical to those derived from actual individual-level data. We demonstrate this approach through real data on 15 068 patient records from 70 clinics at the Children's Hospital of Pennsylvania (CHOP). Assuming that each clinic only shares summary statistics once, we model the COVID-19 PCR test cycle threshold as a function of patient information. Simplicity, communication efficiency, and wider scope of implementation in any statistical software distinguish our approach from existing strategies in the literature.

stat.ME

A Low-Rank Bayesian Approach for Geoadditive Modeling

Kriging is an established methodology for predicting spatial data in geostatistics. Current kriging techniques can handle linear dependencies on spatially referenced covariates. Although splines have shown promise in capturing nonlinear dependencies of covariates, their combination with kriging, especially in handling count data, remains underexplored. This paper proposes a novel Bayesian approach to the low-rank representation of geoadditive models, which integrates splines and kriging to account for both spatial correlations and nonlinear dependencies of covariates. The proposed method accommodates Gaussian and count data inherent in many geospatial datasets. Additionally, Laplace approximations to selected posterior distributions enhances computational efficiency, resulting in faster computation times compared to Markov chain Monte Carlo techniques commonly used for Bayesian inference. Method performance is assessed through a simulation study, demonstrating the effectiveness of the proposed approach. The methodology is applied to the analysis of heavy metal concentrations in the Meuse river and vulnerability to the coronavirus disease 2019 (COVID-19) in Belgium. Through this work, we provide a new flexible and computationally efficient framework for analyzing spatial data.

stat.ME

Inferring age-specific differences in susceptibility to and infectiousness upon SARS-CoV-2 infection based on Belgian social contact data

Several important aspects related to SARS-CoV-2 transmission are not well known due to a lack of appropriate data. However, mathematical and computational tools can be used to extract part of this information from the available data, like some hidden age-related characteristics. In this paper, we present a method to investigate age-specific differences in transmission parameters related to susceptibility to and infectiousness upon contracting SARS-CoV-2 infection. More specifically, we use panel-based social contact data from diary-based surveys conducted in Belgium combined with the next generation principle to infer the relative incidence and we compare this to real-life incidence data. Comparing these two allows for the estimation of age-specific transmission parameters. Our analysis implies the susceptibility in children to be around half of the susceptibility in adults, and even lower for very young children (preschooler). However, the probability of adults and the elderly to contract the infection is decreasing throughout the vaccination campaign, thereby modifying the picture over time.

q-bio.PE

Laplacian-P-splines for Bayesian inference in the mixture cure model

The mixture cure model for analyzing survival data is characterized by the assumption that the population under study is divided into a group of subjects who will experience the event of interest over some finite time horizon and another group of cured subjects who will never experience the event irrespective of the duration of follow-up. When using the Bayesian paradigm for inference in survival models with a cure fraction, it is common practice to rely on Markov chain Monte Carlo (MCMC) methods to sample from posterior distributions. Although computationally feasible, the iterative nature of MCMC often implies long sampling times to explore the target space with chains that may suffer from slow convergence and poor mixing. An alternative strategy for fast and flexible sampling-free Bayesian inference in the mixture cure model is suggested in this paper by combining Laplace approximations and penalized B-splines. A logistic regression model is assumed for the cure proportion and a Cox proportional hazards model with a P-spline approximated baseline hazard is used to specify the conditional survival function of susceptible subjects. Laplace approximations to the conditional latent vector are based on analytical formulas for the gradient and Hessian of the log-likelihood, resulting in a substantial speed-up in approximating posterior distributions. The statistical performance and computational efficiency of the proposed Laplacian-P-splines mixture cure (LPSMC) model is assessed in a simulation study. Results show that LPSMC is an appealing alternative to classic MCMC for approximate Bayesian inference in standard mixture cure models. Finally, the novel LPSMC approach is illustrated on three applications involving real survival data.

stat.ME

Prevalence and trend estimation from observational data with highly variable post-stratification weights

In observational surveys, post-stratification is used to reduce bias resulting from differences between the survey population and the population under investigation. However, this can lead to inflated post-stratification weights and, therefore, appropriate methods are required to obtain less variable estimates. Proposed methods include collapsing post-strata, trimming post-stratification weights, generalized regression estimators (GREG) and weight smoothing models, the latter defined by random-effects models that induce shrinkage across post-stratum means. Here, we first describe the weight-smoothing model for prevalence estimation from binary survey outcomes in observational surveys. Second, we propose an extension of this method for trend estimation. And, third, a method is provided such that the GREG can be used for prevalence and trend estimation for observational surveys. Variance estimates of all methods are described. A simulation study is performed to compare the proposed methods with other established methods. The performance of the nonparametric GREG is consistent over all simulation conditions and therefore serves as a valuable solution for prevalence and trend estimation from observational surveys. The method is applied to the estimation of the prevalence and incidence trend of influenza-like illness using the 2010/2011 Great Influenza Survey in Flanders, Belgium.

stat.AP