SearcharxivSearch

arXiv subjects

Michail Tsagris

Publications and source records attributed to Michail Tsagris.

At least 19 recordsLinked to original sources

Modelling compositional data with structural zero values

Compositional data are positive multivariate data whose sum equals 1. A popular method to analyze such data is via log--ratio transformations, which are however not applicable when zero values are present. In this paper we present a conditional logistic normal distribution suitable for compositional data with structural zero values. The model is applicable to arbitrary dimensions and the EM algorithm guarantees a fast implementation. The regression setting is also presented and a comparison to the Dirichlet analogue is performed.

stat.ME

Non--negative matrix factorization using the \textit{R} package \textsf{nnmf}

Non--negative matrix factorization (NMF) has become an established dimensionality reduction technique for extracting latent structures from non--negative data and has found widespread applications in fields such as bioinformatics, text mining, image analysis, and recommender systems. As the popularity of NMF has increased, numerous \textit{R} packages implementing different optimization strategies and computational frameworks have been developed. Despite their widespread availability, comprehensive evaluations of these implementations under real--world data conditions remain limited. Consequently, researchers often lack objective guidance when selecting an appropriate package for practical applications. This study introduces a new \textit{R} package for NMF and offers asystematic performance comparison with two widely available \textit{R} packages for NMF analysis. Rather than relying on simulated datasets, the evaluation is conducted using real--world data to better reflect the complexity, heterogeneity, and noise characteristics encountered in practical analytical settings. The packages are assessed using a consistent experimental framework, with emphasis on computational efficiency, convergence behavior, reconstruction accuracy, memory utilization, and the stability of the resulting matrix factorization.

stat.ML

A Jensen-Shannon divergence based $k$--$NN$ algorithm for missing value imputation in compositional data

A novel nonparametric method to impute missing values in compositional data is developed. The method is based on the $k$--$NN$ algorithm, utilizes the Jensen-Shannon divergence and employs the Fr{\'e}chet mean to allow for more flexibility in the estimation process. As an extra feature, the hyper-parameters can be self-adaptive according to the pattern of missing values. Unlike restrictive parametric models, the proposed method makes no assumption about the structure of the data and, most importantly, it is applicable even when compositional data contain zero values. Through simulation studies using real data, it is shown that the proposed algorithm outperforms competing algorithms at various settings, not only in terms of accuracy but also in terms of computational efficiency.

stat.ME

Model--based clustering for spherical and hyper--spherical data using elliptically symmetric distributions

Model--based clustering for directional data data has attracted a lot of interest, but most methods utilize rotationally symmetric distributions. This paper suggests the use of elliptically symmetric distributions, namely the elliptically symmetric angular Gaussian and the spherical elliptically symmetric projected Cauchy distributions that were recently proposed in the literature for modelling spherical data. The expectation--maximization algorithm is employed and the inclusion of covariates is also examined. Simulation studies compare the two distributions in terms of choosing the optimal number of clusters and computational cost. We use the mixtures of these two distributions to cluster two datasets on the sphere (earthquake locations) and two hyper--spherical datasets.

stat.ME

On the generalized circular projected Cauchy distribution

\cite{tsagris2025a} proposed the generalized circular projected Cauchy (GCPC) distribution, whose special case is the wrapped Cauchy distribution. In this paper we first derive the relationship with the wrapped Cauchy distribution, and then we attempt to characterize the distribution. We establish the conditions under which the distribution exhibits unimodality. We provide non-analytical formulas for the mean resultant length and the Kullback-Leibler divergence, and analytical form for the cumulative probability function and the entropy of the GCPC distribution. We propose log-likelihood ratio tests for one, or two location parameters without assuming equality of the concentration parameters. We revisit maximum likelihood estimation with and without predictors. In the regression setting we briefly mention the addition of circular and simplicial predictors. Simulation studies illustrate a) the performance of the log-likelihood ratio test when one falsely assumes that the true distribution is the wrapped Cauchy distribution, and b) the empirical rate of convergence of the regression coefficients. Using a real data analysis example we show how to avoid the log-likelihood being trapped in a local maximum and we correct a mistake in the regression setting.

math.ST

Scalable approximation of the transformation-free linear simplicial-simplicial regression via constrained iterative reweighted least squares

Simplicia-simplicial regression concerns statistical modeling scenarios in which both the predictors and the responses are vectors constrained to lie on the simplex. \cite{fiksel2022} introduced a transformation-free linear regression framework for this setting, wherein the regression coefficients are estimated by minimizing the Kullback-Leibler divergence between the observed and fitted compositions, using an expectation-maximization (EM) algorithm for optimization. In this work, we reformulate the problem as a constrained logistic regression model, in line with the methodological perspective of \cite{tsagris2025}, and we obtain parameter estimates via constrained iteratively reweighted least squares. Simulation results indicate that the proposed procedure substantially improves computational efficiency-yielding speed gains ranging from $6\times--326\times$-while providing estimates that closely approximate those obtained from the EM-based approach.

stat.ME

The $\alpha$--regression for compositional data: a unified framework for standard, temporal and spatial regression models including compositional predictors

We revisit the $\alpha$--regression framework for compositional data. We formulate $\alpha$--regression as a non--linear least squares problem, study its asymptotic properties, and provide efficient estimation via the Levenberg--Marquardt algorithm. We then propose a permutation--based hypothesis testing procedure, derive marginal effects for interpretation, and provide a visual inspection of the effect of each predictor. We further discuss robustified versions, the inclusion of natural splines, and the incorporation of compositional predictors, which further facilitate the formulation of a simple time series model. The framework is extended to spatial settings through four models. (a) The $\alpha$--spatially--lagged X regression model, which incorporates spatial spillover effects via spatially--lagged covariates, with decomposition into direct and indirect effects. (b) The $\alpha$--spatial autoregressive model that allows for spatial autocorrelation. (c) The geographically--weighted $\alpha$--regression, which allows coefficients to vary spatially for capturing local relationships. (d) The $\alpha$--eigenvector spatial filtering that is computationally efficient and captures spatial dependence via the eigenvectors of the kernelized distance matrix. Applications to four real datasets illustrate that the models perform on par with or outperform existing models in the literature. The examples showcase that $\alpha$--regression can outperform various competing regression models under different scenarios and its spatial extensions capture the dependence and improve the predictive performance. Overall, the examples provide evidence that the log--ratio methodology does not always lead to the optimal results.

stat.ME

Enhanced Lepage-type test statistics for location-scale shifts with right-skewed data

Detecting simultaneous shifts in location and scale between two populations is a common challenge in statistical inference, particularly in fields like biomedicine where right-skewed data distributions are prevalent. The classical Lepage test, which combines the Wilcoxon-Mann-Whitney and Ansari-Bradley tests, can be suboptimal under these conditions due to its restrictive assumptions of equal variances and medians. This study systematically evaluates enhanced Lepage-type test statistics that incorporate modern robust components for improved performance with right-skewed data. We combine the Fligner-Policello test and Fong-Huang variance estimator for the location component with a novel empirical variance estimator for the Ansari-Bradley scale component, relaxing assumptions of equal variances and medians. Extensive Monte Carlo simulations across exponential, gamma, chi-square, lognormal, and Weibull distributions demonstrate that tests incorporating both robust components achieve power improvements of 10-25\% over the classical Lepage test while maintaining reasonable Type I error control. The practical utility is demonstrated through analyses of four real-world biomedical datasets, where the tests successfully detect significant location-scale shifts. We provide practical guidance for test selection and discuss implementation considerations, making these methods accessible for practitioners in biomedical research and other disciplines where right-skewed data are common.

stat.ME

Simplicial clustering using the $\alpha$--transformation

We introduce two simplicial clustering approaches for compositional data, that are adaptations of the $K$--means and of the Gaussian mixture models algorithms, by employing the $\alpha$--transformation. By utilizing clustering validation indices we can decide on the number of clusters and choose the value of $\alpha$ for the $K$--means, while for the model-based clustering approach information criteria complete this task. extensive simulation studies compare the performance of these two approaches and a real data set illustrates their performance in real world settings.

stat.ME

Fast and light-weight energy statistics using the \textit{R} package \textsf{estats}

Energy statistics ($\mathcal{\varepsilon}$--statistics) enable powerful non-linear dependence measures such as distance correlation, but their computational burden has limited application to large datasets. We present memory-efficient algorithms that compute $\mathcal{\varepsilon}$--statistics related quantities by calculating pairwise distances on-the-fly rather than storing full distance matrices. Our methods achieve 5-156$\times$ speed improvements over existing implementations while reducing memory requirements from $O(n^2)$ to $O(n)$. These advances enable energy statistics computation with sample sizes exceeding tens of thousands observations-previously infeasible with standard implementations-facilitating their use in modern applications across statistics, bioinformatics, and machine learning where large-scale datasets are frequently met. The following cases are demonstrated: energy distance, univariate and multivariate distance variance, distance covariance, (partial) distance correlation and hypothesis testing for the equality of univariate distributions. Functions to compute the aforementioned energy statistics, among others, are available in the \textit{R} package \textsf{estats}.

stat.CO

Energy Based Equality of Distributions Testing for Compositional Data

Not many tests exist for testing the equality for two or more multivariate distributions with compositional data, perhaps due to their constrained sample space. At the moment, there is only one test suggested that relies upon random projections. We propose a novel test termed {\alpha}-Energy Based Test ({\alpha}-EBT) to compare the multivariate distributions of two (or more) compositional data sets. Similar to the aforementioned test, the new test makes no parametric assumptions about the data and, based on simulation studies it exhibits higher power levels.

stat.ME

Circular and Spherical Projected Cauchy Distributions: A Novel Framework for Circular and Directional Data Modeling

We introduce a novel family of projected distributions on the circle and the sphere, namely the circular and spherical projected Cauchy distributions, as promising alternatives for modelling circular and spherical data. The circular distribution encompasses the wrapped Cauchy distribution as a special case, while featuring a more convenient parameterisation. We also propose a generalised wrapped Cauchy distribution that includes an extra parameter, enhancing the fit of the distribution. In the spherical context, we impose two conditions on the scatter matrix of the Cauchy distribution, resulting in an elliptically symmetric distribution. Our projected distributions exhibit attractive properties, such as a closed-form normalising constant and straightforward random value generation. The distribution parameters can be estimated using maximum likelihood, and we assess their bias through numerical studies. Further, we compare our proposed distributions to existing models with real datasets, demonstrating equal or superior fitting both with and without covariates.

stat.ME

Directional data analysis using the spherical Cauchy and the Poisson kernel-based distribution

In 2020, two novel distributions for the analysis of directional data were introduced: the spherical Cauchy distribution and the Poisson kernel-based distribution. This paper provides a detailed exploration of both distributions within various analytical frameworks. To enhance the practical utility of these distributions, alternative parametrizations that offer advantages in numerical stability and parameter estimation are presented, such as implementation of the Newton-Raphson algorithm for parameter estimation, while facilitating a more efficient and simplified approach in the regression framework. Additionally, a two-sample location test based on the log-likelihood ratio test is introduced. This test is designed to assess whether the location parameters of two populations can be assumed equal. The maximum likelihood discriminant analysis framework is developed for classification purposes, and finally, the problem of clustering directional data is addressed, by fitting finite mixtures of Spherical Cauchy or Poisson kernel-based distributions. Empirical validation is conducted through comprehensive simulation studies and real data applications, wherein the performance of the spherical Cauchy and Poisson kernel-based distributions is systematically compared.

stat.ME

Constrained least squares simplicial-simplicial regression

Simplicial-simplicial regression refers to the regression setting where both the responses and predictor variables lie within the simplex space, i.e. they are compositional. For this setting, constrained least squares, where the regression coefficients themselves lie within the simplex, is proposed. The model is transformation-free but the adoption of a power transformation is straightforward, it can treat more than one compositional datasets as predictors and offers the possibility of weights among the simplicial predictors. Among the model's advantages are its ability to treat zeros in a natural way and a highly computationally efficient algorithm to estimate its coefficients. Resampling based hypothesis testing procedures are employed regarding inference, such as linear independence, and equality of the regression coefficients to some pre-specified values. The strategy behind the formulation of the new model is implemented is related to an existing methodology, that is of the same spirit, showcasing how other similar models can be employed as well. Finally, the performance of the proposed technique and its comparison to the existing methodology takes place using simulation studies and real data examples.

stat.ME

Inference for Network Count Time Series with the R Package PNAR

We introduce a new R package useful for inference about network count time series. Such data are frequently encountered in statistics and they are usually treated as multivariate time series. Their statistical analysis is based on linear or log linear models. Nonlinear models, which have been applied successfully in several research areas, have been neglected from such applications mainly because of their computational complexity. We provide R users the flexibility to fit and study nonlinear network count time series models which include either a drift in the intercept or a regime switching mechanism. We develop several computational tools including estimation of various count Network Autoregressive models and fast computational algorithms for testing linearity in standard cases and when non-identifiable parameters hamper the analysis. Finally, we introduce a copula Poisson algorithm for simulating multivariate network count time series. We illustrate the methodology by modeling weekly number of influenza cases in Germany.

stat.ME

Flexible non-parametric regression models for compositional data

Compositional data arise in many real-life applications and versatile methods for properly analyzing this type of data in the regression context are needed. When parametric assumptions do not hold or are difficult to verify, non-parametric regression models can provide a convenient alternative method for prediction. To this end, we consider an extension to the classical $k$--$NN$ regression, termed $α$--$k$--$NN$ regression, that yields a highly flexible non-parametric regression model for compositional data through the use of the $α$-transformation. Unlike many of the recommended regression models for compositional data, zeros values (which commonly occur in practice) are not problematic and they can be incorporated into the proposed models without modification. Extensive simulation studies and real-life data analyses highlight the advantage of using these non-parametric regressions for complex relationships between the compositional response data and Euclidean predictor variables. Both suggest that $α$--$k$--$NN$ regression can lead to more accurate predictions compared to current regression models which assume a, sometimes restrictive, parametric relationship with the predictor variables. In addition, the $α$--$k$--$NN$ regression, in contrast to current regression techniques, enjoys a high computational efficiency rendering it highly attractive for use with large scale, massive, or big data.

stat.ME

Cauchy robust principal component analysis with applications to high-deimensional data sets

Principal component analysis (PCA) is a standard dimensionality reduction technique used in various research and applied fields. From an algorithmic point of view, classical PCA can be formulated in terms of operations on a multivariate Gaussian likelihood. As a consequence of the implied Gaussian formulation, the principal components are not robust to outliers. In this paper, we propose a modified formulation, based on the use of a multivariate Cauchy likelihood instead of the Gaussian likelihood, which has the effect of robustifying the principal components. We present an algorithm to compute these robustified principal components. We additionally derive the relevant influence function of the first component and examine its theoretical properties. Simulation experiments on high-dimensional datasets demonstrate that the estimated principal components based on the Cauchy likelihood outperform or are on par with existing robust PCA techniques.

stat.ME

Modelling structural zeros in compositional data via a zero-censored multivariate normal model

We present a new model for analyzing compositional data with structural zeros. Inspired by \cite{butler2008} who suggested a model in the presence of zero values in the data we propose a model that treats the zero values in a different manner. Instead of projecting every zero value towards a vertex, we project them onto their corresponding edge and fit a zero-censored multivariate model.

stat.ME