SearcharxivSearch

arXiv subjects

Omar Alzeley

Publications and source records attributed to Omar Alzeley.

5 recordsLinked to original sources

Modelling compositional data with structural zero values

Compositional data are positive multivariate data whose sum equals 1. A popular method to analyze such data is via log--ratio transformations, which are however not applicable when zero values are present. In this paper we present a conditional logistic normal distribution suitable for compositional data with structural zero values. The model is applicable to arbitrary dimensions and the EM algorithm guarantees a fast implementation. The regression setting is also presented and a comparison to the Dirichlet analogue is performed.

stat.ME

On the generalized circular projected Cauchy distribution

\cite{tsagris2025a} proposed the generalized circular projected Cauchy (GCPC) distribution, whose special case is the wrapped Cauchy distribution. In this paper we first derive the relationship with the wrapped Cauchy distribution, and then we attempt to characterize the distribution. We establish the conditions under which the distribution exhibits unimodality. We provide non-analytical formulas for the mean resultant length and the Kullback-Leibler divergence, and analytical form for the cumulative probability function and the entropy of the GCPC distribution. We propose log-likelihood ratio tests for one, or two location parameters without assuming equality of the concentration parameters. We revisit maximum likelihood estimation with and without predictors. In the regression setting we briefly mention the addition of circular and simplicial predictors. Simulation studies illustrate a) the performance of the log-likelihood ratio test when one falsely assumes that the true distribution is the wrapped Cauchy distribution, and b) the empirical rate of convergence of the regression coefficients. Using a real data analysis example we show how to avoid the log-likelihood being trapped in a local maximum and we correct a mistake in the regression setting.

math.ST

Scalable approximation of the transformation-free linear simplicial-simplicial regression via constrained iterative reweighted least squares

Simplicia-simplicial regression concerns statistical modeling scenarios in which both the predictors and the responses are vectors constrained to lie on the simplex. \cite{fiksel2022} introduced a transformation-free linear regression framework for this setting, wherein the regression coefficients are estimated by minimizing the Kullback-Leibler divergence between the observed and fitted compositions, using an expectation-maximization (EM) algorithm for optimization. In this work, we reformulate the problem as a constrained logistic regression model, in line with the methodological perspective of \cite{tsagris2025}, and we obtain parameter estimates via constrained iteratively reweighted least squares. Simulation results indicate that the proposed procedure substantially improves computational efficiency-yielding speed gains ranging from $6\times--326\times$-while providing estimates that closely approximate those obtained from the EM-based approach.

stat.ME

Circular and Spherical Projected Cauchy Distributions: A Novel Framework for Circular and Directional Data Modeling

We introduce a novel family of projected distributions on the circle and the sphere, namely the circular and spherical projected Cauchy distributions, as promising alternatives for modelling circular and spherical data. The circular distribution encompasses the wrapped Cauchy distribution as a special case, while featuring a more convenient parameterisation. We also propose a generalised wrapped Cauchy distribution that includes an extra parameter, enhancing the fit of the distribution. In the spherical context, we impose two conditions on the scatter matrix of the Cauchy distribution, resulting in an elliptically symmetric distribution. Our projected distributions exhibit attractive properties, such as a closed-form normalising constant and straightforward random value generation. The distribution parameters can be estimated using maximum likelihood, and we assess their bias through numerical studies. Further, we compare our proposed distributions to existing models with real datasets, demonstrating equal or superior fitting both with and without covariates.

stat.ME

Machine Learning Approach and Extreme Value Theory to Correlated Stochastic Time Series with Application to Tree Ring Data

The main goal of machine learning (ML) is to study and improve mathematical models which can be trained with data provided by the environment to infer the future and to make decisions without necessarily having complete knowledge of all influencing elements. In this work, we describe how ML can be a powerful tool in studying climate modeling. Tree ring growth was used as an implementation in different aspects, for example, studying the history of buildings and environment. By growing and via the time, a new layer of wood to beneath its bark by the tree. After years of growing, time series can be applied via a sequence of tree ring widths. The purpose of this paper is to use ML algorithms and Extreme Value Theory in order to analyse a set of tree ring widths data from nine trees growing in Nottinghamshire. Initially, we start by exploring the data through a variety of descriptive statistical approaches. Transforming data is important at this stage to find out any problem in modelling algorithm. We then use algorithm tuning and ensemble methods to improve the k-nearest neighbors (KNN) algorithm. A comparison between the developed method in this study ad other methods are applied. Also, extreme value of the dataset will be more investigated. The results of the analysis study show that the ML algorithms in the Random Forest method would give accurate results in the analysis of tree ring widths data from nine trees growing in Nottinghamshire with the lowest Root Mean Square Error value. Also, we notice that as the assumed ARMA model parameters increased, the probability of selecting the true model also increased. In terms of the Extreme Value Theory, the Weibull distribution would be a good choice to model tree ring data.

stat.ML