SearcharxivSearch

arXiv subjects

Sean van der Merwe

Publications and source records attributed to Sean van der Merwe.

9 recordsLinked to original sources

A flexible quantile mixed-effects model for censored outcomes

We introduce a Bayesian quantile mixed-effects model for censored longitudinal outcomes based on the skew exponential power (SEP) error distribution. The SEP family separates tail behavior and skewness from the targeted quantile and includes the skew Laplace (SL) distribution as a special case. We derive analytic likelihood contributions for left, right, and interval censoring under the SEP model, so censored observations are handled within a single parametric framework without numerical integration in the likelihood. In simulation studies with varying censoring patterns and skewness profiles, the SEP-based quantile mixed-effects model maintains near-nominal bias and credible interval coverage for regression coefficients. In contrast, the SL-based model can exhibit bias and undercoverage when the data's skewness conflicts with the skewness implied by the target quantile. In an HIV-1 RNA viral load case study with left censoring at the assay limit, bridge-sampled marginal likelihoods and simulation-based residual diagnostics favor the SEP specification across quantiles and yield more stable estimates of treatment-specific viral load trajectories than the SL benchmark.

stat.ME

A robust mixed-effects quantile regression model using generalized Laplace mixtures to handle outliers and skewness

Mixed-effects quantile regression models are widely used to capture heterogeneous responses in hierarchically structured data. The asymmetric Laplace (AL) distribution has traditionally served as the basis for quantile regression; however, its fixed skewness limits flexibility and renders it sensitive to outliers. In contrast, the generalized asymmetric Laplace (GAL) distribution enables more flexible modeling of skewness and heavy-tailed behavior, yet it remains vulnerable to extreme observations. In this paper, we extend the GAL distribution by introducing a contaminated GAL (cGAL) mixture model that incorporates a scale-inflated component to mitigate the impact of outliers without requiring explicit outlier identification or deletion. We apply this model within a Bayesian mixed-effects quantile regression framework to model HIV viral load decay over time. Our results demonstrate that the cGAL-based model more reliably captures the dynamics of HIV viral load decay, yielding more accurate parameter estimates compared to both AL and GAL approaches. Model diagnostics and comparison statistics confirm the cGAL model as the preferred choice. A simulation study further shows that the cGAL model is more robust to outliers than the GAL and exhibits favorable frequentist properties.

stat.ME

Addressing outliers in mixed-effects logistic regression: a more robust modeling approach

This study introduces an outlier-robust model for analyzing hierarchically structured bounded count data within a Bayesian framework, utilizing a logistic regression approach implemented in JAGS. Our model incorporates a t-distributed latent variable to address overdispersion and outliers, improving robustness compared to conventional models such as the beta-binomial, binomial-logit-normal, and standard binomial models. Notably, our model targets a pseudo-median that differs from the true discrete median by less than one count; this closed-form quantity provides a robust and interpretable measure of central tendency. For comparability between all models, we additionally make predictions based on the mean proportion; however, this involves an integration step for the t-distributed nuisance parameter. While limited literature specifically addresses outliers in mixed models for bounded count data, this research fills that gap. The practical utility of the model is demonstrated using a longitudinal medication adherence dataset, where patient behavior often results in abrupt changes and outliers within individual trajectories. A simulation study demonstrates the binomial-logit-t model's strong performance, with comparison statistics favoring it among the four evaluated models. An additional data contamination simulation confirms its robustness against outliers. Our robust approach maintains the integrity of the dataset, effectively handling outliers to provide more accurate and reliable parameter estimates.

stat.ME

On Determining the Distribution of a Goodness-of-Fit Test Statistic

We consider the problem of goodness-of-fit testing for a model that has at least one unknown parameter that cannot be eliminated by transformation. Examples of such problems can be as simple as testing whether a sample consists of independent Gamma observations, or whether a sample consists of independent Generalised Pareto observations given a threshold. Over time the approach to determining the distribution of a test statistic for such a problem has moved towards on-the-fly calculation post observing a sample. Modern approaches include the parametric bootstrap and posterior predictive checks. We argue that these approaches are merely approximations to integrating over the posterior predictive distribution that flows naturally from a given model. Further, we attempt to demonstrate that shortcomings which may be present in the parametric bootstrap, especially in small samples, can be reduced through the use of objective Bayes techniques, in order to more reliably produce a test with the correct size.

stat.ME

Bayesian Fitting of Dirichlet Type I and II Distributions

In his 1986 book, Aitchison explains that compositional data is regularly mishandled in statistical analyses, a pattern that continues to this day. The Dirichlet Type I distribution is a multivariate distribution commonly used to model a set of proportions that sum to one. Aitchinson goes on to lament the difficulties of Dirichlet modelling and the scarcity of alternatives. While he addresses the second of these issues, we address the first. The Dirichlet Type II distribution is a transformation of the Dirichlet Type I distribution and is a multivariate distribution on the positive real numbers with only one more parameter than the number of dimensions. This property of Dirichlet distributions implies advantages over common alternatives as the number of dimensions increase. While not all data is amenable to Dirichlet modelling, there are many cases where the Dirichlet family is the obvious choice. We describe the Dirichlet distributions and show how to fit them using both frequentist and Bayesian methods (we derive and apply two objective priors). The Beta distribution is discussed as a special case. We report a small simulation study to compare the fitting methods. We derive the conditional distributions and posterior predictive conditional distributions. The flexibility of this distribution family is illustrated via examples, the last of which discusses imputation (using the posterior predictive conditional distributions).

math.ST

Bayesian Extreme Value Analysis of Stock Exchange Data

The Solvency II Directive and Solvency Assessment and Management (the South African equivalent) give a Solvency Capital Requirement which is based on a 99.5% Value-at-Risk (VaR) calculation. This calculation involves aggregating individual risks. When considering log returns of financial instruments, especially with share prices, there are extreme losses that are observed from time to time that often do not fit whatever model is proposed for the regular trading behaviour. The problem of accurately modelling these extreme losses is addressed, which, in turn, assists with the calculation of tail probabilities such as the 99.5% VaR. The focus is on the fitting of the Generalized Pareto Distribution (GPD) beyond a threshold. We show how objective Bayes methods can improve parameter estimation and the calculation of risk measures. Lastly we consider the choice of threshold. All aspects are illustrated using share losses on the Johannesburg Stock Exchange (JSE).

stat.AP

Time Series Analysis of the Southern Oscillation Index using Bayesian Additive Regression Trees

Bayesian additive regression trees (BART) is a regression technique developed by Chipman et al. (2008). Its usefulness in standard regression settings has been clearly demonstrated, but it has not been applied to time series analysis as yet. We discuss the difficulties in applying this technique to time series analysis and demonstrate its superior predictive capabilities in the case of a well know time series: the Southern Oscillation Index.

stat.AP

A method for Bayesian regression modelling of composition data

Many scientific and industrial processes produce data that is best analysed as vectors of relative values, often called compositions or proportions. The Dirichlet distribution is a natural distribution to use for composition or proportion data. It has the advantage of a low number of parameters, making it the parsimonious choice in many cases. In this paper we consider the case where the outcome of a process is Dirichlet, dependent on one or more explanatory variables in a regression setting. We explore some existing approaches to this problem, and then introduce a new simulation approach to fitting such models, based on the Bayesian framework. We illustrate the advantages of the new approach through simulated examples and an application in sport science. These advantages include: increased accuracy of fit, increased power for inference, and the ability to introduce random effects without additional complexity in the analysis.

stat.ME

An empirical study to order citation statistics between subject fields

An empirical study is conducted to compare citations per publication, statistics and observed Hirsch indexes between subject fields using summary statistics of countries. No distributional assumptions are made and ratios are calculated. These ratios can be used to make approximate comparisons between researchers of different subject fields with respect to the Hirsch index.

cs.DL