SearcharxivSearch

arXiv subjects

Francesca Greselin

Publications and source records attributed to Francesca Greselin.

13 recordsLinked to original sources

Robust fuzzy clustering with cellwise outliers

In a data matrix, we may distinguish between cases, each represented by a row vector for a statistical unit, and cells, which correspond to single entries of the data matrix. Recent developments in Robust Statistics have introduced the cellwise contamination paradigm, which assumes contamination on cells rather than on entire cases. This approach becomes particularly relevant as the number of variables increases. Indeed, discarding or downweighting entire cases because of a few anomalous cells in them, as done by traditional (casewise) robust methods, can result in substantial information loss, since the non-contaminated (or reliable) cells can still be highly informative. This philosophy can also be considered in fuzzy clustering, by assuming that reliable cells within a case may still provide useful information for determining fuzzy memberships. A robust fuzzy clustering proposal is thus introduced in this work, combining the advantages of dealing with outlying cells and simultaneously controlling the degree of fuzziness of unit assignments. The cluster-specific relationships among variables, detected by the fuzzy clustering approach, are also key to better identifying outlying cells and correct them. The strengths of the proposed methodology are illustrated through a simulation study and two real-world applications. The effects of the model's tuning parameters are explored, and some guidance for users on how to set them suitably is provided.

stat.ME

Cellwise outlier detection in heterogeneous populations

Real-world applications may be affected by outlying values. In the model-based clustering literature, several methodologies have been proposed to detect units that deviate from the majority of the data (rowwise outliers) and trim them from the parameter estimates. However, the discarded observations can encompass valuable information in some observed features. Following the more recent cellwise contamination paradigm, we introduce a Gaussian mixture model for cellwise outlier detection. The proposal is estimated via an Expectation-Maximization (EM) algorithm with an additional step for flagging the contaminated cells of a data matrix and then imputing - instead of discarding - them before the parameter estimation. This procedure adheres to the spirit of the EM algorithm by treating the contaminated cells as missing values. We analyze the performance of the proposed model in comparison with other existing methodologies through a simulation study with different scenarios and illustrate its potential use for clustering, outlier detection, and imputation on three real data sets. Additional applications include socio-economic studies, environmental analysis, healthcare, and any domain where the aim is to cluster data affected by missing information and outlying values within features.

stat.ME

Measuring income inequality via percentile relativities

"The rich are getting richer" implies that the population income distributions are getting more right skewed and heavily tailed. For such distributions, the mean is not the best measure of the center, but the classical indices of income inequality, including the celebrated Gini index, are all mean-based. In view of this, Professor Gastwirth sounded an alarm back in 2014 by suggesting to incorporate the median into the definition of the Gini index, although noted a few shortcomings of his proposed index. In the present paper we make a further step in the modification of classical indices and, to acknowledge the possibility of differing viewpoints, arrive at three median-based indices of inequality. They avoid the shortcomings of the previous indices and can be used even when populations are ultra heavily tailed, that is, when their first moments are infinite. The new indices are illustrated both analytically and numerically using parametric families of income distributions, and further illustrated using capital incomes coming from 2001 and 2018 surveys of fifteen European countries. We also discuss the performance of the indices from the perspective of income transfers.

stat.ME

A text analysis for Operational Risk loss descriptions

Financial institutions manage operational risk (OpRisk) by carrying out activities required by regulation, such as collecting loss data, calculating capital requirements, and reporting. For this purpose, for each OpRisk event, loss amounts, dates, organizational units involved, event types, and descriptions are recorded in the OpRisk databases. In recent years, operational risk functions have been required to go beyond their regulatory tasks to proactively manage operational risk, preventing or mitigating its impact. As OpRisk databases also contain event descriptions, an area of opportunity is to extract information from such texts. The present work introduces for the first time a structured workflow for the application of text analysis techniques (one of the main Natural Language Processing tasks) to the OpRisk event descriptions to identify managerial clusters (more granular than regulatory categories) representing the root-causes of the underlying risks. We have complemented and enriched the established framework of statistical methods based on quantitative data. Specifically, after delicate tasks like data cleaning, text vectorization, and semantic adjustment, we have applied methods of dimensionality reduction and several clustering models with algorithms to compare their performances and weaknesses. Our results improve retrospective knowledge of loss events and enable to mitigate future risks.

stat.AP

A Two-Stage Bayesian Semiparametric Model for Novelty Detection with Robust Prior Information

Novelty detection methods aim at partitioning the test units into already observed and previously unseen patterns. However, two significant issues arise: there may be considerable interest in identifying specific structures within the novelty, and contamination in the known classes could completely blur the actual separation between manifest and new groups. Motivated by these problems, we propose a two-stage Bayesian semiparametric novelty detector, building upon prior information robustly extracted from a set of complete learning units. We devise a general-purpose multivariate methodology that we also extend to handle functional data objects. We provide insights on the model behavior by investigating the theoretical properties of the associated semiparametric prior. From the computational point of view, we propose a suitable $\boldsymbolξ$-sequence to construct an independent slice-efficient sampler that takes into account the difference between manifest and novelty components. We showcase our model performance through an extensive simulation study and applications on both multivariate and functional datasets, in which diverse and distinctive unknown patterns are discovered.

stat.AP

Robust variable selection in the framework of classification with label noise and outliers: applications to spectroscopic data in agri-food

Classification of high-dimensional spectroscopic data is a common task in analytical chemistry. Well-established procedures like support vector machines (SVMs) and partial least squares discriminant analysis (PLS-DA) are the most common methods for tackling this supervised learning problem. Nonetheless, interpretation of these models remains sometimes difficult, and solutions based on feature selection are often adopted as they lead to the automatic identification of the most informative wavelengths. Unfortunately, for some delicate applications like food authenticity, mislabeled and adulterated spectra occur both in the calibration and/or validation sets, with dramatic effects on the model development, its prediction accuracy and robustness. Motivated by these issues, the present paper proposes a robust model-based method that simultaneously performs variable selection, outliers and label noise detection. We demonstrate the effectiveness of our proposal in dealing with three agri-food spectroscopic studies, where several forms of perturbations are considered. Our approach succeeds in diminishing problem complexity, identifying anomalous spectra and attaining competitive predictive accuracy considering a very low number of selected wavelengths.

stat.AP

Robust variable selection for model-based learning in presence of adulteration

The problem of identifying the most discriminating features when performing supervised learning has been extensively investigated. In particular, several methods for variable selection in model-based classification have been proposed. Surprisingly, the impact of outliers and wrongly labeled units on the determination of relevant predictors has received far less attention, with almost no dedicated methodologies available in the literature. In the present paper, we introduce two robust variable selection approaches: one that embeds a robust classifier within a greedy-forward selection procedure and the other based on the theory of maximum likelihood estimation and irrelevance. The former recasts the feature identification as a model selection problem, while the latter regards the relevant subset as a model parameter to be estimated. The benefits of the proposed methods, in contrast with non-robust solutions, are assessed via an experiment on synthetic data. An application to a high-dimensional classification problem of contaminated spectroscopic data concludes the paper.

stat.AP

The Social Welfare Implications of the Zenga Index

We introduce the social welfare implications of the Zenga index, a recently proposed index of inequality. Our proposal is derived by following the seminal book by Son (2011) and the recent working paper by Kakwani and Son (2019). We compare the Zenga based approach with the classical one, based on the Lorenz curve and the Gini coefficient, as well as the Bonferroni index. We show that the social welfare specification based on the Zenga uniformity curve presents some peculiarities that distinguish it from the other considered indexes. The social welfare specification presented here provides a deeper understanding of how the Zenga index evaluates the inequality in a distribution.

econ.GN

Anomaly and Novelty detection for robust semi-supervised learning

Three important issues are often encountered in Supervised and Semi-Supervised Classification: class-memberships are unreliable for some training units (label noise), a proportion of observations might depart from the main structure of the data (outliers) and new groups in the test set may have not been encountered earlier in the learning phase (unobserved classes). The present work introduces a robust and adaptive Discriminant Analysis rule, capable of handling situations in which one or more of the afore-mentioned problems occur. Two EM-based classifiers are proposed: the first one that jointly exploits the training and test sets (transductive approach), and the second one that expands the parameter estimate using the test set, to complete the group structure learned from the training set (inductive approach). Experiments on synthetic and real data, artificially adulterated, are provided to underline the benefits of the proposed method.

stat.AP

A robust approach to model-based classification based on trimming and constraints

In a standard classification framework a set of trustworthy learning data are employed to build a decision rule, with the final aim of classifying unlabelled units belonging to the test set. Therefore, unreliable labelled observations, namely outliers and data with incorrect labels, can strongly undermine the classifier performance, especially if the training size is small. The present work introduces a robust modification to the Model-Based Classification framework, employing impartial trimming and constraints on the ratio between the maximum and the minimum eigenvalue of the group scatter matrices. The proposed method effectively handles noise presence in both response and exploratory variables, providing reliable classification even when dealing with contaminated datasets. A robust information criterion is proposed for model selection. Experiments on real and simulated data, artificially adulterated, are provided to underline the benefits of the proposed method.

stat.AP

Inferential results for a new measure of inequality

In this paper we derive inferential results for a new index of inequality, specifically defined for capturing significant changes observed both in the left and in the right tail of the income distributions. The latter shifts are an apparent fact for many countries like US, Germany, UK, and France in the last decades, and are a concern for many policy makers. We propose two empirical estimators for the index, and show that they are asymptotically equivalent. Afterwards, we adopt one estimator and prove its consistency and asymptotic normality. Finally we introduce an empirical estimator for its variance and provide conditions to show its convergence to the finite theoretical value. An analysis of real data on net income from the Bank of Italy Survey of Income and Wealth is also presented, on the base of the obtained inferential results.

math.ST

Measuring economic inequality and risk: a unifying approach based on personal gambles, societal preferences and references

The underlying idea behind the construction of indices of economic inequality is based on measuring deviations of various portions of low incomes from certain references or benchmarks, that could be point measures like population mean or median, or curves like the hypotenuse of the right triangle where every Lorenz curve falls into. In this paper we argue that by appropriately choosing population-based references, called societal references, and distributions of personal positions, called gambles, which are random, we can meaningfully unify classical and contemporary indices of economic inequality, as well as various measures of risk. To illustrate the herein proposed approach, we put forward and explore a risk measure that takes into account the relativity of large risks with respect to small ones.

stat.ME

Maximum likelihood estimation in constrained parameter spaces for mixtures of factor analyzers

Mixtures of factor analyzers are becoming more and more popular in the area of model based clustering of high-dimensional data. According to the likelihood approach in data modeling, it is well known that the unconstrained log-likelihood function may present spurious maxima and singularities and this is due to specific patterns of the estimated covariance structure, when their determinant approaches 0. To reduce such drawbacks, in this paper we introduce a procedure for the parameter estimation of mixtures of factor analyzers, which maximizes the likelihood function in a constrained parameter space. We then analyze and measure its performance, compared to the usual non-constrained approach, via some simulations and applications to real data sets.

stat.ME