SearcharxivSearch

arXiv subjects

Carmela Iorio

Publications and source records attributed to Carmela Iorio.

5 recordsLinked to original sources

A Family of Divergence Measures for Evaluating the Reconstruction Quality of Explainable Ensemble Trees

Validating interpretable surrogate models for ensemble learners requires measuring agreement between the ensemble's internal representation and its surrogate approximation, rather than mere association. Correlation-based approaches are scale-invariant and fail to detect systematic discrepancies in co-occurrence structure. We propose a statistical framework grounded in the agreement-association distinction, centered on the normalized Loss of Interpretability (nLoI). Rooted in the Cressie-Read power divergence family with lambda equal to 2, the nLoI admits a closed-form decomposition into within-node and between-node components, providing a unique diagnostic capability to identify precisely where and why reconstruction fails. The framework incorporates four complementary measures capturing distinct structural facets of approximation quality. A unified permutation testing procedure delivers valid inference for all measures within a single resampling pass. Theoretical properties, including boundedness and symmetry, are established for each metric. Monte Carlo simulations and empirical evaluations confirm exact Type I error control and demonstrate that these measures detect reconstruction fidelity gradients invisible to correlation-based alternatives. The framework is developed and illustrated in the context of Explainable Ensemble Trees (E2Tree), and empirical evaluation on three benchmark datasets illustrates the practical utility of the framework.

cs.LG

Extending Explainable Ensemble Trees (E2Tree) to regression contexts

Ensemble methods such as random forests have transformed the landscape of supervised learning, offering highly accurate prediction through the aggregation of multiple weak learners. However, despite their effectiveness, these methods often lack transparency, impeding users' comprehension of how RF models arrive at their predictions. Explainable ensemble trees (E2Tree) is a novel methodology for explaining random forests, that provides a graphical representation of the relationship between response variables and predictors. A striking characteristic of E2Tree is that it not only accounts for the effects of predictor variables on the response but also accounts for associations between the predictor variables through the computation and use of dissimilarity measures. The E2Tree methodology was initially proposed for use in classification tasks. In this paper, we extend the methodology to encompass regression contexts. To demonstrate the explanatory power of the proposed algorithm, we illustrate its use on real-world datasets.

cs.LG

Adjusted Concordance Index, an extension of the Adjusted Rand index to fuzzy partitions

In comparing clustering partitions, Rand index (RI) and Adjusted Rand index (ARI) are commonly used for measuring the agreement between the partitions. Both these external validation indexes aim to analyze how close is a cluster to a reference (or to prior knowledge about the data) by counting corrected classified pairs of elements. When the aim is to evaluate the solution of a fuzzy clustering algorithm, the computation of these measures require converting the soft partitions into hard ones. It is known that different fuzzy partitions describing very different structures in the data can lead to the same crisp partition and consequently to the same values of these measures. We compare the existing approaches to evaluate the external validation criteria in fuzzy clustering and we propose an extension of the ARI for fuzzy partitions based on the normalized degree of concordance. Through use of real and simulated data, we analyze and evaluate the performance of our proposal.

stat.ME

Boosted-Oriented Probabilistic Smoothing-Spline Clustering of Series

Fuzzy clustering methods allow the objects to belong to several clusters simultaneously, with different degrees of membership. However, a factor that influences the performance of fuzzy algorithms is the value of fuzzifier parameter. In this paper, we propose a fuzzy clustering procedure for data (time) series that does not depend on the definition of a fuzzifier parameter. It comes from two approaches, theoretically motivated for unsupervised and supervised classification cases, respectively. The first is the Probabilistic Distance (PD) clustering procedure. The second is the well known Boosting philosophy. Our idea is to adopt a boosting prospective for unsupervised learning problems, in particular we face with non hierarchical clustering problems. The aim is to assign each instance (i.e. a series) of a data set to a cluster. We assume the representative instance of a given cluster (i.e. the cluster center) as a target instance, a loss function as a synthetic index of the global performance and the probability of each instance to belong to a given cluster as the individual contribution of a given instance to the overall solution. The global performance of the proposed method is investigated by various experiments.

stat.ME

Parsimonious Time Series Clustering

We introduce a parsimonious model-based framework for clustering time course data. In these applications the computational burden becomes often an issue due to the number of available observations. The measured time series can also be very noisy and sparse and a suitable model describing them can be hard to define. We propose to model the observed measurements by using P-spline smoothers and to cluster the functional objects as summarized by the optimal spline coefficients. In principle, this idea can be adopted within all the most common clustering frameworks. In this work we discuss applications based on a k-means algorithm. We evaluate the accuracy and the efficiency of our proposal by simulations and by dealing with drosophila melanogaster gene expression data.

stat.ME