SearcharxivSearch

arXiv subjects

Ricardo Maronna

Publications and source records attributed to Ricardo Maronna.

4 recordsLinked to original sources

Robust Model-Based Clustering

We propose a new class of robust and Fisher-consistent estimators for mixture models. These estimators can be used to construct robust model-based clustering procedures. We study in detail the case of multivariate normal mixtures and propose a procedure that uses S estimators of multivariate location and scatter. We develop an algorithm to compute the estimators and to build the clusters which is quite similar to the EM algorithm. An extensive Monte Carlo simulation study shows that our proposal compares favorably with other robust and non robust model-based clustering procedures. We apply ours and alternative procedures to a real data set and again find that the best results are obtained using our proposal.

stat.ME

Robust multivariate methods in Chemometrics

This chapter presents an introduction to robust statistics with applications of a chemometric nature. Following a description of the basic ideas and concepts behind robust statistics, including how robust estimators can be conceived, the chapter builds up to the construction (and use) of robust alternatives for some methods for multivariate analysis frequently used in chemometrics, such as principal component analysis and partial least squares. The chapter then provides an insight into how these robust methods can be used or extended to classification. To conclude, the issue of validation of the results is being addressed: it is shown how uncertainty statements associated with robust estimates, can be obtained.

stat.ME

Improving the Peña-Prieto "KSD" procedure

Peña and Prieto (2007) proposed the "Kurtosis plus specific directions" (KSD) method for robust multivariate location and scatter estimation and outlier detection. Maronna and Yohai (2017) employed it as an initial estimator for multivariate S- and MM-estimators, and their simulations showed that KSD generally outperforms initial estimators based on subsampling. However further simulations show that KSD may become unstable and give wrong results in extreme situations when the contamination rate is "high" (>=0.2) and the ratio n/p of cases to variables is "low" (<10). Two simple modifications of the procedure are proposed, which greatly improve on the method's performance as an initial estimator, with only a small increase in computational time.

stat.ME

High finite-sample efficiency and robustness based on distance-constrained maximum likelihood

Good robust estimators can be tuned to combine a high breakdown point and a specified asymptotic efficiency at a central model. This happens in regression with MM- and tau-estimators among others. However, the finite-sample efficiency of these estimators can be much lower than the asymptotic one. To overcome this drawback, an approach is proposed for parametric models, which is based on a distance between parameters. Given a robust estimator, the proposed one is obtained by maximizing the likelihood under the constraint that the distance is less than a given threshold. For the linear model with normal errors and using the MM estimator and the distance induced by the Kullback-Leibler divergence, simulations show that the proposed estimator attains a finite-sample efficiency close to one, while its maximum mean squared error under pointwise outlier contamination is smaller than that of the MM estimator. The same approach also shows good results in the estimation of multivariate location and scatter.

math.ST