SearcharxivSearch

arXiv subjects

Maurizio Vichi

Publications and source records attributed to Maurizio Vichi.

5 recordsLinked to original sources

Nonparametric framework for the definition, adaptive detection and probabilistic interpretation of outliers

Outlier detection is a fundamental challenge in data processing, with critical implications for robustness across statistical modeling, machine learning and exploratory data analysis. However, existing proposals rarely offer a universal, domain-agnostic definition of an outlier, often relying on heuristic trimming quotas that lack a statistical interpretation. To address this, we propose a nonparametric framework built on a pseudo-isolation outlier score. This score enables a formal, probabilistic definition of an anomaly tied to a false-alarm rate $α$, which extends into a rigorous geometric classification of internal and external outliers. We show that this mechanism seamlessly embeds into any objective-based clustering framework to identify cluster-specific outliers. Here, we integrate it into $K$-means to create ODK-means. The inferential capabilities and topological properties of this framework are explored both theoretically through formal propositions and empirically through extensive simulations and methodological tutorials, highlighting the practical actionability and intuitive appeal of the proposed outlier detection logic.

stat.ME

Bi-dendrograms for clustering the categories of a multivariate categorical data set

The clustering of categories in a multivariate categorical data set is investigated, where the problem separates into that of merging categories of the same variables (i.e., within-variable categories), and combining categories of different variables (i.e., between-variable categories). For the within-variable problem, the objective is to arrive at fewer categories (and, consequently, lower data dimensionality) without affecting the essential features of the data set, thereby simplifying the interpretation of any analysis using the categorical variables. The categories can be of an ordinal or nominal nature, and this property is respected in the clustering, where only adjacent categories of ordinal variables can be combined. For the between-variable problem, the objective is to arrive at asmall number of category clusters that typify the observations in the data set. In this latter problem there is no restriction on which categories can combine, as long as they do not combine within the same variable. In each of these problems, results are given in the form of a pair of dendrograms stacked one on top of the other, called a bi-dendrogram. For the within-variable problem, once all categories within each variable have been merged, the second stage is to cluster the variables themselves. For the between-variable problem, the second stage is to cluster groups of respondents that fall into the response sets arrived at in the first stage of clustering. The approach is illustrated using a sociological survey data set from the International Social Survey Program.

stat.ME

Spherical Double K-Means: a co-clustering approach for text data analysis

In text analysis, Spherical K-means (SKM) is a specialized k-means clustering algorithm widely utilized for grouping documents represented in high-dimensional, sparse term-document matrices, often normalized using techniques like TF-IDF. Researchers frequently seek to cluster not only documents but also the terms associated with them into coherent groups. To address this dual clustering requirement, we introduce Spherical Double K-Means (SDKM), a novel methodology that simultaneously clusters documents and terms. This approach offers several advantages: first, by integrating the clustering of documents and terms, SDKM provides deeper insights into the relationships between content and vocabulary, enabling more effective topic identification and keyword extraction. Additionally, the two-level clustering assists in understanding both overarching themes and specific terminologies within document clusters, enhancing interpretability. SDKM effectively handles the high dimensionality and sparsity inherent in text data by utilizing cosine similarity, leading to improved computational efficiency. Moreover, the method captures dynamic changes in thematic content over time, making it well-suited for applications in rapidly evolving fields. Ultimately, SDKM presents a comprehensive framework for advancing text mining efforts, facilitating the uncovering of nuanced patterns and structures that are critical for robust data analysis. We apply SDKM to the corpus of US presidential inaugural addresses, spanning from George Washington in 1789 to Joe Biden in 2021. Our analysis reveals distinct clusters of words and documents that correspond to significant historical themes and periods, showcasing the method's ability to facilitate a deeper understanding of the data. Our findings demonstrate the efficacy of SDKM in uncovering underlying patterns in textual data.

stat.ME

Structural Equation Modeling and simultaneous clustering through the Partial Least Squares algorithm

The identification of different homogeneous groups of observations and their appropriate analysis in PLS-SEM has become a critical issue in many appli- cation fields. Usually, both SEM and PLS-SEM assume the homogeneity of all units on which the model is estimated, and approaches of segmentation present in literature, consist in estimating separate models for each segments of statistical units, which have been obtained either by assigning the units to segments a priori defined. However, these approaches are not fully accept- able because no causal structure among the variables is postulated. In other words, a modeling approach should be used, where the obtained clusters are homogeneous with respect to the structural causal relationships. In this paper, a new methodology for simultaneous non-hierarchical clus- tering and PLS-SEM is proposed. This methodology is motivated by the fact that the sequential approach of applying first SEM or PLS-SEM and second the clustering algorithm such as K-means on the latent scores of the SEM/PLS-SEM may fail to find the correct clustering structure existing in the data. A simulation study and an application on real data are included to evaluate the performance of the proposed methodology.

stat.ME

Time-varying clustering of multivariate longitudinal observations

We propose a statistical method for clustering of multivariate longitudinal data into homogeneous groups. This method relies on a time-varying extension on the classical K-means algorithm, where a multivariate vector autoregressive model is additionally assumed for modeling the evolution of clusters' centroids over time. We base the inference on a least squares specification of the model and coordinate descent algorithm. To illustrate our work, we consider a longitudinal dataset on human development. Three variables are modeled, namely life expectancy, education and gross domestic product.

stat.ME