Searcharxiv⌕ Search

arXiv subjects

Janet van Niekerk

Publications and source records attributed to Janet van Niekerk.

At least 19 recordsLinked to original sources

Informative Distance-Based Priors for Correlation Matrices Centred on a Target Reference

Specifying a prior over the space of correlation matrices is a persistent challenge in Bayesian analysis. The space is a curved manifold whose dimension grows quadratically with the number of variables, making substantive prior beliefs difficult to encode.\\ We propose a distance-based prior that assigns mass decaying exponentially in the Fisher arc-length distance from a user-specified reference correlation matrix, enabling shrinkage toward any target correlation structure rather than being confined to the identity matrix. Formally, this is constructed as a Penalised Complexity prior, but its interpretation shifts accordingly: unless the chosen target represents a structurally simpler state, the shrinkage penalises deviation rather than complexity in the usual sense. To accommodate conditional independence constraints, we introduce a parameterisation that constructs the correlation matrix via the Cholesky factor of the inverse correlation matrix with respect to a user-supplied graph, thereby reducing the number of free parameters from one per variable pair to one per graph edge. The prior is proper for every positive value of its rate parameter, accommodates correlations of either sign under any graph structure, and reduces to a fully unstructured prior when the graph is complete. A direct sampling algorithm is provided, enabling prior predictive checks and sensitivity analysis, implemented within the \texttt{graphpcor} package.

stat.ME↗

A flexible wrapped Lindley-type distribution for angular data modelling

Flexible distributions for modelling angular data have received considerable attention in recent years, with ongoing work extending existing circular models to provide greater flexibility in capturing diverse angular behaviours. In this paper, we introduce and study the w3PL distribution, a circular model obtained by extending the wrapped Lindley distribution by incorporating two additional shape parameters. The proposed generalisation increases flexibility in modelling concentration and skewness while preserving analytical tractability and encompassing existing circular models as special cases. Closed-form expressions for the probability density function, cumulative distribution function, and trigonometric moments are derived, allowing key distributional properties to be studied analytically. The distributional modality is characterised, and the nature of invariance is investigated for the newly proposed circular model. Parameter estimation is developed within a regularised maximum likelihood framework, and a simulation study demonstrates reliable parameter recovery and stable finite-sample performance. Applications to angular datasets from geology, marine biology, and finance illustrate the model's practical significance and show improved fit relative to existing circular alternatives.

stat.ME↗

A robust contaminated discrete Weibull regression model for outlier-prone count data

Count data often exhibit overdispersion driven by heavy tails or excess zeros, making standard models (e.g., Poisson, negative binomial) insufficient for handling outlying observations. We propose a novel contaminated discrete Weibull (cDW) framework that augments a baseline discrete Weibull (DW) distribution with a heavier-tail subcomponent. This mixture retains a single shifted-median parameter for a unified regression link while selectively assigning extreme outcomes to the heavier-tail subdistribution. The cDW distribution accommodates strictly positive data by setting the truncation limit c=1 as well as full-range counts with c=0. We develop a Bayesian regression formulation and describe posterior inference using Markov chain Monte Carlo sampling. In an application to hospital length-of-stay data (with c=1, meaning the minimum possible stay is 1), the cDW model more effectively captures extreme stays and preserves the median-based link. Simulation-based residual checks, leave-one-out cross-validation, and a Kullback-Leibler outlier assessment confirm that the cDW model provides a more robust fit than the single-component DW model, reducing the influence of outliers and improving predictive accuracy. A simulation study further demonstrates the cDW model's robustness in the presence of heavy contamination. We also discuss how a hurdle scheme can accommodate datasets with many zeros while preventing the spurious inflation of zeros in situations without genuine zero inflation.

stat.ME↗

Outlier-robust copula regression for bivariate continuous proportions: an application to cushion plant vitality

Continuous proportions measured on the same experimental unit often pose two challenges: interior outliers that inflate variance beyond the beta ceiling and residual dependence that invalidates independent-margin models. We introduce a Bayesian copula modeling approach that combines rectangular-beta margins, which temper interior outliers by reallocating mass from the peak to a uniform component, with a single-parameter copula to capture concordance. Gaussian, Gumbel, and Clayton copula families are fitted, and log marginal likelihoods are obtained via bridge sampling to guide model selection. Applied to a 13-year survey (2003-2016) of Azorella selago cushion plants on sub-Antarctic Marion Island, the copula models outperform independence baselines in explaining percent dead stem cover. Accounting for between-year dependence uncovers a positive west-slope effect and weakens the cushion size effect. Simulation results show negligible bias and near-nominal 95% highest posterior density coverage across a range of tail weight and dependence scenarios, confirming good frequentist properties. The method integrates readily with JAGS and provides a robust default for paired proportion data in ecology and other disciplines where bounded outcomes and occasional outliers coincide.

stat.ME↗

Spatial Gaussian fields for complex areas with application to marine megafauna conservation

Spatial Gaussian fields (SGFs) are widely employed in modeling the distributions of marine megafauna, yet they traditionally rely on assumptions of isotropy and stationarity, conditions that often prove unrealistic in complex ecological environments featuring coastlines, islands, and depth gradients acting as partial movement barriers. Existing spatial models typically treat these barriers as either fully impermeable, completely blocking species movement and dispersal, or entirely absent, which inadequately represents most real-world scenarios. To address this limitation, we introduce the Transparent Barrier Model, an extension of spatial Gaussian fields that explicitly incorporates barriers with varying levels of permeability. The model assigns spatially varying range parameters to distinct barrier regions, allowing ecological and geographical knowledge about barrier permeability to directly inform model specifications. This approach maintains computational efficiency by utilizing the integrated nested Laplace approximation (INLA) framework combined with stochastic partial differential equations (SPDEs), ensuring feasible application even in large, complex spatial domains.We demonstrate the practical utility and flexibility of the Transparent Barrier Model through its application to dugong (Dugong dugon) distribution data from the Red Sea.

stat.ME↗

A graphical framework for interpretable correlation matrix models

In this work, we present a new approach for constructing models for correlation matrices with a user-defined graphical structure. The graphical structure makes correlation matrices interpretable and avoids the quadratic increase of parameters as a function of the dimension. We suggest an automatic approach to define a prior using a natural sequence of simpler models within the Penalized Complexity framework for the unknown parameters in these models. We illustrate this approach with three applications: a multivariate linear regression of four biomarkers, a multivariate disease mapping, and a multivariate longitudinal joint modelling. Each application underscores our method's intuitive appeal, signifying a substantial advancement toward a more cohesive and enlightening model that facilitates a meaningful interpretation of correlation matrices.

stat.ME↗

Scalable skewed Bayesian inference for latent Gaussian models

Approximate Bayesian inference for the class of latent Gaussian models can be achieved efficiently with integrated nested Laplace approximations (INLA). Based on recent reformulations in the INLA methodology, we propose a further extension that is necessary in some cases like heavy-tailed likelihoods or binary regression with imbalanced data. This extension formulates a skewed version of the Laplace method such that some marginals are skewed and some are kept Gaussian while the dependence is maintained with the Gaussian copula from the Laplace method. Our approach is formulated to be scalable in model and data size, using a variational inferential framework enveloped in INLA. We illustrate the necessity and performance using simulated cases, as well as a case study of a rare disease where class imbalance is naturally present.

stat.ME↗

Joint Modeling of Multivariate Longitudinal and Survival Outcomes with the R package INLAjoint

This paper introduces the R package INLAjoint, designed as a toolbox for fitting a diverse range of regression models addressing both longitudinal and survival outcomes. INLAjoint relies on the computational efficiency of the integrated nested Laplace approximations methodology, an efficient alternative to Markov chain Monte Carlo for Bayesian inference, ensuring both speed and accuracy in parameter estimation and uncertainty quantification. The package facilitates the construction of complex joint models by treating individual regression models as building blocks, which can be assembled to address specific research questions. Joint models are relevant in biomedical studies where the collection of longitudinal markers alongside censored survival times is common. They have gained significant interest in recent literature, demonstrating the ability to rectify biases present in separate modeling approaches such as informative censoring by a survival event or confusion bias due to population heterogeneity. We provide a comprehensive overview of the joint modeling framework embedded in INLAjoint with illustrative examples. Through these examples, we demonstrate the practical utility of INLAjoint in handling complex data scenarios encountered in biomedical research.

stat.ME↗

Bayesian survival analysis with INLA

This tutorial shows how various Bayesian survival models can be fitted using the integrated nested Laplace approximation in a clear, legible, and comprehensible manner using the INLA and INLAjoint R-packages. Such models include accelerated failure time, proportional hazards, mixture cure, competing risks, multi-state, frailty, and joint models of longitudinal and survival data, originally presented in the article "Bayesian survival analysis with BUGS" (Alvares et al., 2021). In addition, we illustrate the implementation of a new joint model for a longitudinal semicontinuous marker, recurrent events, and a terminal event. Our proposal aims to provide the reader with syntax examples for implementing survival models using a fast and accurate approximate Bayesian inferential approach.

stat.ME↗

Spatio-temporal insights for wind energy harvesting in South Africa

Understanding complex spatial dependency structures is a crucial consideration when attempting to build a modeling framework for wind speeds. Ideally, wind speed modeling should be very efficient since the wind speed can vary significantly from day to day or even hour to hour. But complex models usually require high computational resources. This paper illustrates how to construct and implement a hierarchical Bayesian model for wind speeds using the Weibull density function based on a continuously-indexed spatial field. For efficient (near real-time) inference the proposed model is implemented in the r package R-INLA, based on the integrated nested Laplace approximation (INLA). Specific attention is given to the theoretical and practical considerations of including a spatial component within a Bayesian hierarchical model. The proposed model is then applied and evaluated using a large volume of real data sourced from the coastal regions of South Africa between 2011 and 2021. By projecting the mean and standard deviation of the Matern field, the results show that the spatial modeling component is effectively capturing variation in wind speeds which cannot be explained by the other model components. The mean of the spatial field varies between $\pm 0.3$ across the domain. These insights are valuable for planning and implementation of green energy resources such as wind farms in South Africa. Furthermore, shortcomings in the spatial sampling domain is evident in the analysis and this is important for future sampling strategies. The proposed model, and the conglomerated dataset, can serve as a foundational framework for future investigations into wind energy in South Africa.

stat.ME↗

Low-rank variational Bayes correction to the Laplace method

Approximate inference methods like the Laplace method, Laplace approximations and variational methods, amongst others, are popular methods when exact inference is not feasible due to the complexity of the model or the abundance of data. In this paper we propose a hybrid approximate method called Low-Rank Variational Bayes correction (VBC), that uses the Laplace method and subsequently a Variational Bayes correction in a lower dimension, to the joint posterior mean. The cost is essentially that of the Laplace method which ensures scalability of the method, in both model complexity and data size. Models with fixed and unknown hyperparameters are considered, for simulated and real examples, for small and large datasets.

stat.ME↗

Fast and flexible inference for joint models of multivariate longitudinal and survival data using Integrated Nested Laplace Approximations

Modeling longitudinal and survival data jointly offers many advantages such as addressing measurement error and missing data in the longitudinal processes, understanding and quantifying the association between the longitudinal markers and the survival events and predicting the risk of events based on the longitudinal markers. A joint model involves multiple submodels (one for each longitudinal/survival outcome) usually linked together through correlated or shared random effects. Their estimation is computationally expensive (particularly due to a multidimensional integration of the likelihood over the random effects distribution) so that inference methods become rapidly intractable, and restricts applications of joint models to a small number of longitudinal markers and/or random effects. We introduce a Bayesian approximation based on the Integrated Nested Laplace Approximation algorithm implemented in the R package R-INLA to alleviate the computational burden and allow the estimation of multivariate joint models with fewer restrictions. Our simulation studies show that R-INLA substantially reduces the computation time and the variability of the parameter estimates compared to alternative estimation strategies. We further apply the methodology to analyze 5 longitudinal markers (3 continuous, 1 count, 1 binary, and 16 random effects) and competing risks of death and transplantation in a clinical trial on primary biliary cholangitis. R-INLA provides a fast and reliable inference technique for applying joint models to the complex multivariate data encountered in health research.

stat.ME↗

Non-stationary Bayesian Spatial Model for Disease Mapping based on Sub-regions

This paper aims to extend the Besag model, a widely used Bayesian spatial model in disease mapping, to a non-stationary spatial model for irregular lattice-type data. The goal is to improve the model's ability to capture complex spatial dependence patterns and increase interpretability. The proposed model uses multiple precision parameters, accounting for different intensities of spatial dependence in different sub-regions. We derive a joint penalized complexity prior for the flexible local precision parameters to prevent overfitting and ensure contraction to the stationary model at a user-defined rate. The proposed methodology can be used as a basis for the development of various other non-stationary effects over other domains such as time. An accompanying R package 'fbesag' equips the reader with the necessary tools for immediate use and application. We illustrate the novelty of the proposal by modeling the risk of dengue in Brazil, where the stationary spatial assumption fails and interesting risk profiles are estimated when accounting for spatial non-stationary.

stat.ME↗

Bayesian Estimation of Two-Part Joint Models for a Longitudinal Semicontinuous Biomarker and a Terminal Event with R-INLA: Interests for Cancer Clinical Trial Evaluation

Two-part joint models for a longitudinal semicontinuous biomarker and a terminal event have been recently introduced based on frequentist estimation. The biomarker distribution is decomposed into a probability of positive value and the expected value among positive values. Shared random effects can represent the association structure between the biomarker and the terminal event. The computational burden increases compared to standard joint models with a single regression model for the biomarker. In this context, the frequentist estimation implemented in the R package frailtypack can be challenging for complex models (i.e., large number of parameters and dimension of the random effects). As an alternative, we propose a Bayesian estimation of two-part joint models based on the Integrated Nested Laplace Approximation (INLA) algorithm to alleviate the computational burden and fit more complex models. Our simulation studies confirm that INLA provides accurate approximation of posterior estimates and to reduced computation time and variability of estimates compared to frailtypack in the situations considered. We contrast the Bayesian and frequentist approaches in the analysis of two randomized cancer clinical trials (GERCOR and PRIME studies), where INLA has a reduced variability for the association between the biomarker and the risk of event. Moreover, the Bayesian approach was able to characterize subgroups of patients associated with different responses to treatment in the PRIME study. Our study suggests that the Bayesian approach using INLA algorithm enables to fit complex joint models that might be of interest in a wide range of clinical applications.

stat.ME↗

A new avenue for Bayesian inference with INLA

Integrated Nested Laplace Approximations (INLA) has been a successful approximate Bayesian inference framework since its proposal by Rue et al. (2009). The increased computational efficiency and accuracy when compared with sampling-based methods for Bayesian inference like MCMC methods, are some contributors to its success. Ongoing research in the INLA methodology and implementation thereof in the R package R-INLA, ensures continued relevance for practitioners and improved performance and applicability of INLA. The era of big data and some recent research developments, presents an opportunity to reformulate some aspects of the classic INLA formulation, to achieve even faster inference, improved numerical stability and scalability. The improvement is especially noticeable for data-rich models. We demonstrate the efficiency gains with various examples of data-rich models, like Cox's proportional hazards model, an item-response theory model, a spatial model including prediction, and a 3-dimensional model for fMRI data.

stat.ME↗

Parallelized integrated nested Laplace approximations for fast Bayesian inference

There is a growing demand for performing larger-scale Bayesian inference tasks, arising from greater data availability and higher-dimensional model parameter spaces. In this work we present parallelization strategies for the methodology of integrated nested Laplace approximations (INLA), a popular framework for performing approximate Bayesian inference on the class of Latent Gaussian models. Our approach makes use of nested OpenMP parallelism, a parallel line search procedure using robust regression in INLA's optimization phase and the state-of-the-art sparse linear solver PARDISO. We leverage mutually independent function evaluations in the algorithm as well as advanced sparse linear algebra techniques. This way we can flexibly utilize the power of today's multi-core architectures. We demonstrate the performance of our new parallelization scheme on a number of different real-world applications. The introduction of parallelism leads to speedups of a factor 10 and more for all larger models. Our work is already integrated in the current version of the open-source R-INLA package, making its improved performance conveniently available to all users.

stat.CO↗

An Extended Simplified Laplace strategy for Approximate Bayesian inference of Latent Gaussian Models using R-INLA

Various computational challenges arise when applying Bayesian inference approaches to complex hierarchical models. Sampling-based inference methods, such as Markov Chain Monte Carlo strategies, are renowned for providing accurate results but with high computational costs and slow or questionable convergence. On the contrary, approximate methods like the Integrated Nested Laplace Approximation (INLA) construct a deterministic approximation to the univariate posteriors through nested Laplace Approximations. This method enables fast inference performance in Latent Gaussian Models, which encode a large class of hierarchical models. R-INLA software mainly consists of three strategies to compute all the required posterior approximations depending on the accuracy requirements. The Simplified Laplace approximation (SLA) is the most attractive because of its speed performance since it is based on a Taylor expansion up to order three of a full Laplace Approximation. Here we enhance the methodology by simplifying the computations necessary for the skewness and modal configuration. Then we propose an expansion up to order four and use the Extended Skew Normal distribution as a new parametric fit. The resulting approximations to the marginal posterior densities are more accurate than those calculated with the SLA, with essentially no additional cost.

stat.ME↗

Joint Quantile Disease Mapping with Application to Malaria and G6PD Deficiency

Statistical analysis based on quantile regression methods is more comprehensive, flexible, and less sensitive to outliers when compared to mean regression methods. When the link between different diseases are of interest, joint disease mapping is useful for measuring directional correlation between them. Most studies study this link through multiple correlated mean regressions. In this paper we propose a joint quantile regression framework for multiple diseases where different quantile levels can be considered. We are motivated by the theorized link between the presence of Malaria and the gene deficiency G6PD, where medical scientist have anecdotally discovered a possible link between high levels of G6PD and lower than expected levels of Malaria initially pointing towards the occurrence of G6PD inhibiting the occurrence of Malaria. This link cannot be investigated with mean regressions and thus the need for flexible joint quantile regression in a disease mapping framework. Our joint quantile disease mapping model can be used for linear and non-linear effects of covariates by stochastic splines, since we define it as a latent Gaussian model. We perform Bayesian inference of this model using the INLA framework embedded in the R software package INLA. Finally, we illustrate the applicability of model by analyzing the malaria and G6PD deficiency incidences in 21 African countries using linked quantiles of different levels.

stat.ME↗