SearcharxivSearch

arXiv subjects

Arno Otto

Publications and source records attributed to Arno Otto.

3 recordsLinked to original sources

Modelling and detecting mild and gross anomalies in circular data via double-contaminated models

In this paper, we propose a model-based framework to robustify inference for circular data in the presence of anomalous observations, distinguishing between mild and gross anomalies. Starting from a unimodal and symmetric reference model on $[0,2\pi)$, parametrized by a mean direction and concentration, we construct a family of finite mixtures: a gross-anomaly model obtained by adding a circular uniform component; a mild-anomaly (contaminated) model obtained by mixing the reference distribution with a less concentrated version sharing the same mean direction; and a general three-component specification combining both models, the double-contaminated model. Posterior component probabilities provide an automatic classification of observations without ad hoc thresholds, while the mixing weights yield interpretable measures of anomaly prevalence and dispersion inflation. For illustration, we consider two classical circular reference distributions, the wrapped normal and von Mises. The methodology is evaluated through an extensive simulation study and three real-data applications involving animal movement directions and wind directions. The results indicate that jointly modelling mild and gross departures improves model fit and yields an informative decomposition of the directional data, demonstrating that mixture-based robustness is valuable not only for anomaly detection but also for the interpretation and the identification of latent structure in directional data.

stat.ME

A Contaminated Model for Overdispersed Multinomial Microbiome Count Data

Multinomial count data, such as microbial composition profiles derived from sequencing studies, frequently contain anomalous observations that distort parameter estimates. The Dirichlet-multinomial (DM) distribution is widely used in this setting but remains sensitive to such contamination. We propose the contaminated Dirichlet-multinomial (CDM) distribution, a two-component mixture in which the regular data come from a DM component with a lower dispersion and the irregular data come from a DM component with an inflated dispersion parameter. This construction accommodates anomalies without requiring their removal, and yields a natural rule for anomaly detection via posterior probabilities. Through sensitivity analyses involving both single-point anomalies and background noise, we demonstrate that the CDM distribution effectively downweights the influence of anomalous observations on the parameter estimates. The model is applied to gut microbiome data from a colorectal carcinogenesis study, where it consistently outperforms the DM distribution across all information criteria and identifies biologically plausible anomaly proportions in both the healthy and carcinoma subsets.

stat.ME

Mean regression for (0,1) responses via beta scale mixtures

To achieve a greater general flexibility for modeling heavy-tailed bounded responses, a beta scale mixture model is proposed. Each member of the family is obtained by multiplying the scale parameter of the conditional beta distribution by a mixing random variable taking values on all or part of the positive real line and whose distribution depends on a single parameter governing the tail behavior of the resulting compound distribution. These family members allow for a wider range of values for skewness and kurtosis. To validate the effectiveness of the proposed model, we conduct experiments on both simulated data and real datasets. The results indicate that the beta scale mixture model demonstrates superior performance relative to the classical beta regression model and alternative competing methods for modeling responses on the bounded unit domain.

stat.ME