SearcharxivSearch

arXiv subjects

Andriette Bekker

Publications and source records attributed to Andriette Bekker.

16 recordsLinked to original sources

Handling Missingness and Censoring in Dirichlet Mixture Models

Incomplete compositional data analysis faces a fundamental limitation: likelihood-based methods for compositional models generally require fully observed compositions, making it difficult to accommodate missing or censored proportions directly on the simplex. Consequently, analysts often discard partially observed compositions or transform the data into unconstrained spaces, potentially sacrificing interpretability and coherence. This paper proposes a likelihood-based method for incomplete compositional data without leaving the simplex. Specifically, we develop an Expectation-Maximisation (EM) type algorithm for fitting finite mixtures of Dirichlet distributions in the presence of missing and censored components. The proposed approach performs parameter estimation and model-based imputation simultaneously while preserving the compositional structure and interpretability of the original variables. A simulation experiment evaluates the performance of the proposed estimators and imputations under increasingly complex coarsening mechanisms. Particular attention is paid to clustering performance, and model selection outcomes. The results showed beneficial clustering performance despite observations being incomplete, and a higher probability of model selection metrics identifying the correct number of clusters compared to current alternative of case-deletion. The practical utility of the method is illustrated using two real datasets with distinct coarsened patterns. Analysis of the xenolith dataset identifies a four-component Dirichlet mixture that reveals interpretable profiles of rock types and speciation methods. Application to PM$_{2.5}$ speciation data from the Air Quality System, containing both left-censored and missing-at-random values, supports a four-component mixture model that characterises compositional parts of particulate matter across the United States.

stat.ME

Sleep pattern profiling using a finite mixture of contaminated multivariate skew-normal distributions on incomplete data

Medical data often exhibit characteristics that make cluster analysis particularly challenging, such as missing values, outliers, and cluster features like skewness. Typically, such data would need to be preprocessed -- by cleaning outliers and missing values -- before clustering could be performed. However, these preliminary steps rely on objective functions different from those used in the clustering stage. In this paper, we propose a unified model-based clustering approach that simultaneously handles atypical observations, missing values, and cluster-wise skewness within a single framework. Each cluster is modelled using a contaminated multivariate skew-normal distribution -- a convenient two-component mixture of multivariate skew-normal densities -- in which one component represents the main data (the "bulk") and the other captures potential outliers. From an inferential perspective, we implement and use a variant of the EM algorithm to obtain the maximum likelihood estimates of the model parameters. Simulation studies demonstrate that the proposed model outperforms existing approaches in both clustering accuracy and outlier detection, across low- and high-dimensional settings, even in the presence of substantial missingness. The method is further applied to the Cleveland Children's Sleep and Health Study (CCSHS), a dataset characterised by incomplete observations. Without any preprocessing, the proposed approach identifies five distinct groups of sleepers, revealing meaningful differences in sleeper typologies.

stat.ME

Clustering data with values missing at random using scale mixtures of multivariate skew-normal distributions

Handling missing data is a major challenge in model-based clustering, especially when the data exhibit skewness and heavy tails. We address this by extending the finite mixture of scale mixtures of multivariate skew-normal (FMSMSN) family to accommodate incomplete data under a missing at random (MAR) mechanism. Unlike previous work that is limited to one of the special cases of the FMSMSN family, our method offers a cluster analysis methodology for the entire family that accounts for skewness and excess kurtosis amidst data with missing values. The multivariate skew-normal distribution, as parameterised by \cite{azzalini1996} and \cite{arnoldbeaver} includes the normal distribution as a special case, which ensures that our method is flexible toward existing symmetric model-based clustering techniques under a normality assumption. We derive the distributional properties of the missing components of the data and propose an augmented EM-type algorithm tailored for incomplete observations. The modified E-step yields closed-form expressions for the conditional expectations of the missing values. The simulation experiments showcase the flexibility of the FMSMSN family in both clustering performance and parameter recovery for varying percentages of missing values, while incorporating the effects of sample size and cluster proximity. Finally, we illustrate the practical utility of the proposed method by applying special cases of the FMSMSN family to global CO2 emissions data.

stat.ME

Bayesian Semi-Parametric Spatial Dispersed Count Model for Precipitation Analysis

The appropriateness of the Poisson model is frequently challenged when examining spatial count data marked by unbalanced distributions, over-dispersion, or under-dispersion. Moreover, traditional parametric models may inadequately capture the relationships among variables when covariates display ambiguous functional forms or when spatial patterns are intricate and indeterminate. To tackle these issues, we propose an innovative Bayesian hierarchical modeling system. This method combines non-parametric techniques with an adapted dispersed count model based on renewal theory, facilitating the effective management of unequal dispersion, non-linear correlations, and complex geographic dependencies in count data. We illustrate the efficacy of our strategy by applying it to lung and bronchus cancer mortality data from Iowa, emphasizing environmental and demographic factors like ozone concentrations, PM2.5, green space, and asthma prevalence. Our analysis demonstrates considerable regional heterogeneity and non-linear relationships, providing important insights into the impact of environmental and health-related factors on cancer death rates. This application highlights the significance of our methodology in public health research, where precise modeling and forecasting are essential for guiding policy and intervention efforts. Additionally, we performed a simulation study to assess the resilience and accuracy of the suggested method, validating its superiority in managing dispersion and capturing intricate spatial patterns relative to conventional methods. The suggested framework presents a flexible and robust instrument for geographical count analysis, offering innovative insights for academics and practitioners in disciplines such as epidemiology, environmental science, and spatial statistics.

stat.ME

Does wind affect the orientation of vegetation stripes? A copula-based mixture model for axial and circular data

Motivated by a case study of vegetation patterns, we introduce a mixture model with concomitant variables to examine the association between the orientation of vegetation stripes and wind direction. The proposal relies on a novel copula-based bivariate distribution for mixed axial and circular observations and provides a parsimonious and computationally tractable approach to examine the dependence of two environmental variables observed in a complex manifold. The findings suggest that dominant winds shape the orientation of vegetation stripes through a mechanism of neighbouring plants providing wind shelter to downwind individuals.

stat.AP

Neutrosophic Birnbaum-Saunders distribution with applications

Classical statistics deals with determined and precise data analysis. But in reality, there are many cases where the information is not accurate and a degree of impreciseness, uncertainty, incompleteness, and vagueness is observed. In these situations, uncertainties can make classical statistics less accurate. That is where neutrosophic statistics steps in to improve accuracy in data analysis. In this article, we consider the Birnbaum-Saunders distribution (BSD) which is very flexible and practical for real world data modeling. By integrating the neutrosophic concept, we improve the BSD's ability to manage uncertainty effectively. In addition, we provide maximum likelihood parameter estimates. Subsequently, we illustrate the practical advantages of the neutrosophic model using two cases from the industrial and environmental fields. This paper emphasizes the significance of the neutrosophic BSD as a robust solution for modeling and analysing imprecise data, filling a crucial gap left by classical statistical methods.

stat.AP

Type I multivariate Pólya-Aeppli distributions with applications

An extensive body of literature exists that specifically addresses the univariate case of zero-inflated count models. In contrast, research pertaining to multivariate models is notably less developed. We proposed two new parsimonious multivariate models which can be used to model correlated multivariate overdispersed count data. Furthermore, for different parameter settings and sample sizes, various simulations are performed. In conclusion, we demonstrated the performance of the newly proposed multivariate candidates on two benchmark datasets, which surpasses that of several alternative approaches.

stat.ME

A dependent circular-linear model for multivariate biomechanical data: Ilizarov ring fixator study

Biomechanical and orthopaedic studies frequently encounter complex datasets that encompass both circular and linear variables. In most cases the circular and linear variables are (i) considered in isolation with dependency between variables neglected and (ii) the cyclicity of the circular variables disregarded resulting in erroneous decision making. Given the inherent characteristics of circular variables, it is imperative to adopt methods that integrate directional statistics to achieve precise modelling. This paper is motivated by the modelling of biomechanical data, i.e., the fracture displacements, that is used as a measure in external fixator comparisons. We focus on a data set, based on an Ilizarov ring fixator, comprising of six variables. A modelling framework applicable to the 6D joint distribution of circular-linear data based on vine copulas is proposed. The pair-copula decomposition concept of vine copulas represents the dependence structure as a combination of circular-linear, circular-circular and linear-linear pairs modelled by their respective copulas. This framework allows us to assess the dependencies in the joint distribution as well as account for the cyclicity of the circular variables. Thus, a new approach for accurate modelling of mechanical behaviour for Ilizarov ring fixators and other data of this nature is imparted.

stat.AP

A spatial analysis of COVID-19 reported cases in the Gauteng province, South Africa: Identifying wards to be targeted early in future infectious diseases outbreak

The COVID-19 pandemic caused major disruptions and contributed to the loss of livelihoods and income. The pandemic also provided public health and health systems policy shifts towards better promotion and protection in responding to such disasters and emergencies. Due to differing effects of socio-economic infectious disease vulnerabilities and pre-pandemic levels of preparedness for health emergencies, health system strengthening requires targeted and ununiform implementation. We employ spatial statistical methods on the COVID-19 confirmed cases in identifying wards that could be targeted for strengthening health security in the Gauteng Province, South Africa. In this way, the identified high-risk wards would be more effective and prepared to respond to future pandemics and emergencies.

stat.AP

In search of the perfect fit: interpretation, flexible modelling, and the existing generalisations of the normal distribution

Many generalised distributions exist for modelling data with vastly diverse characteristics. However, very few of these generalisations of the normal distribution have shape parameters with clear roles that determine, for instance, skewness and tail shape. In this chapter, we review existing skewing mechanisms and their properties in detail. Using the knowledge acquired, we add a skewness parameter to the body-tail generalised normal distribution \cite{BTGN}, that yields the \ac{FIN} with parameters for location, scale, body-shape, skewness, and tail weight. Basic statistical properties of the \ac{FIN} are provided, such as the \ac{PDF}, cumulative distribution function, moments, and likelihood equations. Additionally, the \ac{FIN} \ac{PDF} is extended to a multivariate setting using a student t-copula, yielding the \ac{MFIN}. The \ac{MFIN} is applied to stock returns data, where it outperforms the t-copula multivariate generalised hyperbolic, Azzalini skew-t, hyperbolic, and normal inverse Gaussian distributions.

stat.ME

Spatio-temporal insights for wind energy harvesting in South Africa

Understanding complex spatial dependency structures is a crucial consideration when attempting to build a modeling framework for wind speeds. Ideally, wind speed modeling should be very efficient since the wind speed can vary significantly from day to day or even hour to hour. But complex models usually require high computational resources. This paper illustrates how to construct and implement a hierarchical Bayesian model for wind speeds using the Weibull density function based on a continuously-indexed spatial field. For efficient (near real-time) inference the proposed model is implemented in the r package R-INLA, based on the integrated nested Laplace approximation (INLA). Specific attention is given to the theoretical and practical considerations of including a spatial component within a Bayesian hierarchical model. The proposed model is then applied and evaluated using a large volume of real data sourced from the coastal regions of South Africa between 2011 and 2021. By projecting the mean and standard deviation of the Matern field, the results show that the spatial modeling component is effectively capturing variation in wind speeds which cannot be explained by the other model components. The mean of the spatial field varies between $\pm 0.3$ across the domain. These insights are valuable for planning and implementation of green energy resources such as wind farms in South Africa. Furthermore, shortcomings in the spatial sampling domain is evident in the analysis and this is important for future sampling strategies. The proposed model, and the conglomerated dataset, can serve as a foundational framework for future investigations into wind energy in South Africa.

stat.ME

Directional Gaussian spatial processes for South African wind data

Accurate wind pattern modelling is crucial for various applications, including renewable energy, agriculture, and climate adaptation. In this paper, we introduce the wrapped Gaussian spatial process (WGSP), as well as the projected Gaussian spatial process (PGSP) custom-tailored for South Africa's intricate wind behaviour. Unlike conventional models struggling with the circular nature of wind direction, the WGSP and PGSP adeptly incorporate circular statistics to address this challenge. Leveraging historical data sourced from meteorological stations throughout South Africa, the WGSP and PGSP significantly increase predictive accuracy while capturing the nuanced spatial dependencies inherent to wind patterns. The superiority of the PGSP model in capturing the structural characteristics of the South African wind data is evident. As opposed to the PGSP, the WGSP model is computationally less demanding, allows for the use of less informative priors, and its parameters are more easily interpretable. The implications of this study are far-reaching, offering potential benefits ranging from the optimisation of renewable energy systems to the informed decision-making in agriculture and climate adaptation strategies. The WGSP and PGSP emerge as robust and invaluable tools, facilitating precise modelling of wind patterns within the dynamic context of South Africa.

stat.AP

Uncovering a generalised gamma distribution: from shape to interpretation

In this paper, we introduce the flexible interpretable gamma (FIG) distribution which has been derived by Weibullisation of the body-tail generalised normal distribution. The parameters of the FIG have been verified graphically and mathematically as having interpretable roles in controlling the left-tail, body, and right-tail shape. The generalised gamma (GG) distribution has become a staple model for positive data in statistics due to its interpretable parameters and tractable equations. Although there are many generalised forms of the GG which can provide better fit to data, none of them extend the GG so that the parameters are interpretable. Additionally, we present some mathematical characteristics and prove the identifiability of the FIG parameters. Finally, we apply the FIG model to hand grip strength and insurance loss data to assess its flexibility relative to existing models.

math.ST

Protein Structure Parameterization via Mobius Distributions on the Torus

Proteins constitute a large group of macromolecules with a multitude of functions for all living organisms. Proteins achieve this by adopting distinct three-dimensional structures encoded by the sequence of their constituent amino acids in one or more polypeptides. In this paper, the statistical modelling of the protein backbone torsion angles is considered. Two new distributions are proposed for toroidal data by applying the Möbius transformation to the bivariate von Mises distribution. Marginal and conditional distributions in addition to sine-skewed versions of the proposed models are also developed. Three big data sets consisting of bivariate information about protein domains are analysed to illustrate the strength of the flexible proposed models. Finally, a simulation study is done to evaluate the obtained maximum likelihood estimates and also to find the best method of generating samples from the proposed models to use as the proposal distributions in the Markov Chain Monte Carlo sampling method for predicting the 3D structure of proteins.

stat.ME

Empowering Differential Networks Using Bayesian Analysis

Differential networks (DN) are important tools for modeling the changes in conditional dependencies between multiple samples. A Bayesian approach for estimating DNs, from the classical viewpoint, is introduced with a computationally efficient threshold selection for graphical model determination. The algorithm separately estimates the precision matrices of the DN using the Bayesian adaptive graphical lasso procedure. Synthetic experiments illustrate that the Bayesian DN performs exceptionally well in numerical accuracy and graphical structure determination in comparison to state-of-the-art methods. The proposed method is applied to South African COVID-$19$ data to investigate the change in DN structure between various phases of the pandemic.

stat.ME

Spatial analysis and prediction of COVID-19 spread in South Africa after lockdown

What is the impact of COVID-19 on South Africa? This paper envisages assisting researchers and decision-makers in battling the COVID-19 pandemic focusing on South Africa. This paper focuses on the spread of the disease by applying heatmap retrieval of hotspot areas and spatial analysis is carried out using the Moran index. For capturing spatial autocorrelation between the provinces of South Africa, the adjacent, as well as the geographical distance measures, are used as a weight matrix for both absolute and relative counts. Furthermore, generalized logistic growth curve modeling is used for the prediction of the COVID-19 spread. We expect this data-driven modeling to provide some insights into hotspot identification and timeous action controlling the spread of the virus.

physics.soc-ph