SearcharxivSearch

arXiv subjects

Claudio Agostinelli

Publications and source records attributed to Claudio Agostinelli.

At least 19 recordsLinked to original sources

A Composite Divergence Approach to Robust Multivariate Estimation under Cellwise and Casewise Contamination

Composite likelihood (CL) methods provide a computationally efficient alternative to full likelihood inference for complex multivariate models by replacing the joint likelihood with a product of lower-dimensional marginal or conditional components. Like the MLE, however, the maximum CL estimator (MCLE) is highly sensitive to data contamination. On the other hand, robust divergence-based procedures such as the minimum density power divergence (DPD) estimator require the full joint density and so scale poorly to complex multivariate models. We introduce the composite DPD (CDPD), a genuine statistical divergence built entirely from the low-dimensional component densities defining a CL, combining the computational scalability of CL with the robustness of the DPD. The resulting minimum CDPD estimator (MCDPDE) robustifies the MCLE without requiring integration over the full multivariate sample space. We establish consistency, asymptotic normality, and the influence function of the MCDPDE under regularity conditions on the component models alone, without requiring correct specification of the full joint distribution. We show that it is qualitatively robust for every positive value of its tuning parameter, unlike the MCLE recovered as the limit. Because its components can be chosen at the pairwise or cell level, the framework guards simultaneously against casewise and cellwise contamination. Operating directly on component densities rather than elliptical distance structures, it extends robust inference beyond the elliptical models to which most existing cellwise-robust procedures are confined. We develop computational algorithms implemented in the accompanying R package mvdpd. Simulation studies and real-data applications show that the MCDPDE achieves substantial robustness gains over the MCLE while retaining competitive efficiency under the assumed model.

math.ST

Bayesian Variable Selection in Generalized Linear Models

Covariate selection in Generalized Linear Models (GLMs) is a fundamental problem in statistics, as including irrelevant predictors might lead to overfitting and poor interpretability, while omitting relevant ones might result in biased estimates. Most Bayesian approaches to variable selection -- including spike-and-slab priors and continuous shrinkage priors -- have key limitations, e.g., (i) are based on non fully conjugate formulations, (ii) are restricted to a linear model, or (iii) lack posterior consistency guarantees for the variable selection procedure and model parameters. In this work, we propose a fully Bayesian hierarchical and conjugate framework for covariate selection in GLMs, applicable to any distribution in the exponential family, based on modeling a binary inclusion indicator that directly encodes covariate inclusion in the linear predictor. In our approach, variable selection and parameter estimation are performed simultaneously, incorporating both sources of uncertainty in posterior inference. Consequently, our methodology provides a valid post-model Bayesian selection procedure. We present theoretical guarantees of the proposed fully conjugate Bayesian variable selection for GLMs, establishing posterior consistency of both the inclusion indicators and the active regression coefficients. We derive an efficient Gibbs Sampling algorithm with a corresponding R package implementation. We validate the proposed method on synthetic and real-world datasets, demonstrating competitive predictive and inferential performance.

stat.ME

Comments on "Challenges of cellwise outliers" by Jakob Raymaekers and Peter J. Rousseeuw

The main aim of robust statistics is the development of methods able to cope with the presence of outliers. A new type of outliers, namely "cellwise", has garnered considerable attention. The state of the art for dealing with cellwise contamination in different models is presented in Raymaekers and Rousseeuw (2024). Outliers in time series can be treated as cellwise outliers, a further discussion on this subject is presented.

stat.ME

Central subspace data depth

Statistical data depth plays an important role in the analysis of multivariate data sets. The main outcome is a center-outward ordering of the observations that can be used both to highlight features of the underlying distribution of the data and as input to further statistical analysis. An important property of data depth is related to symmetric distributions as the point with the highest depth value, the center, coincides with the point of symmetry. However, there are applications in which it is more natural to consider symmetry with respect to a subspace of a certain dimension rather than to a point, i.e. a subspace of dimension zero. We provide a general framework to construct statistical data depths which attain maximum value in a subspace, providing a center-outward ordering from that subspace. We refer to these data depths as central subspace data depths. Moreover, if the distribution is symmetric with respect to a subspace, then the depth is maximized at that subspace. We introduce general notions of symmetry about a subspace for distributions, study the properties of central subspace data depths and provide asymptotic convergence for the corresponding sample versions. Additionally, we discuss connections with projection pursuit and dimension reduction. An application based on custom data fraud detection shows the importance of the proposed approach and strengthens its potential.

math.ST

Hellinger loss function for Generative Adversarial Networks

We propose Hellinger-type loss functions for training Generative Adversarial Networks (GANs), motivated by the boundedness, symmetry, and robustness properties of the Hellinger distance. We define an adversarial objective based on this divergence and study its statistical properties within a general parametric framework. We establish the existence, uniqueness, consistency, and joint asymptotic normality of the estimators obtained from the adversarial training procedure. In particular, we analyze the joint estimation of both generator and discriminator parameters, offering a comprehensive asymptotic characterization of the resulting estimators. We introduce two implementations of the Hellinger-type loss and we evaluate their empirical behavior in comparison with the classic (Maximum Likelihood-type) GAN loss. Through a controlled simulation study, we demonstrate that both proposed losses yield improved estimation accuracy and robustness under increasing levels of data contamination.

stat.ML

A Weighted Likelihood Approach Based on Statistical Data Depths

We propose a general approach to construct weighted likelihood estimating equations with the aim of obtaining robust parameter estimates. We modify the standard likelihood equations by incorporating a weight that reflects the statistical depth of each data point relative to the model, as opposed to the sample. An observation is considered regular when the corresponding difference of these two depths is close to zero. When this difference is large the observation score contribution is downweighted. We study the asymptotic properties of the proposed estimator, including consistency and asymptotic normality, for a broad class of weight functions. In particular, we establish asymptotic normality under the standard regularity conditions typically assumed for the maximum likelihood estimator (MLE). Our weighted likelihood estimator achieves the same asymptotic efficiency as the MLE in the absence of contamination, while maintaining a high degree of robustness in contaminated settings. In stark contrast to the traditional minimum divergence/disparity estimators, our results hold even if the dimension of the data diverges with the sample size, without requiring additional assumptions on the existence or smoothness of the underlying densities. We also derive the finite sample breakdown point of our estimator for both location and scatter matrix in the elliptically symmetric model. Detailed results and examples are presented for robust parameter estimation in the multivariate normal model. Robustness is further illustrated using two real data sets and a Monte Carlo simulation study.

math.ST

Vector-Valued Gaussian Processes and their Kernels on a Class of Metric Graphs

Despite the increasing importance of stochastic processes on linear networks and graphs, current literature on multivariate (vector-valued) Gaussian random fields on metric graphs is elusive. This paper challenges several aspects related to the construction of proper matrix-valued kernels structures. We start by considering matrix-valued metrics that can be composed with scalar- or matrix-valued functions to implement valid kernels associated with vector-valued Gaussian fields. We then provide conditions for certain classes of matrix-valued functions to be composed with the univariate resistance metric and ensure positive semidefiniteness. Special attention is then devoted to Euclidean trees, where a substantial effort is required given the absence of literature related to multivariate kernels depending on the $\ell_1$ metric. Hence, we provide a foundational contribution to certain classes of matrix-valued positive semidefinite functions depending on the $\ell_1$ metric. This fact is then used to characterise kernels on Euclidean trees with a finite number of leaves. Amongst those, we provide classes of matrix-valued covariance functions that are compactly supported.

math.ST

Estimation of a multivariate von Mises distribution for contaminated torus data

The occurrence of atypical circular observations on the torus can badly affect parameter estimation of the multivariate von Mises distribution. This paper addresses the problem of robust fitting of the multivariate von Mises model using the weighted likelihood methodology. The key ingredients are non-parametric density estimation for multivariate circular data and the definition of appropriate weighted estimating equations. Some theoretical properties are discussed. The finite sample behavior of the proposed weighted likelihood estimator has been investigated by Monte Carlo numerical studies and empirical applications.

stat.ME

Weighted likelihood methods for robust fitting of wrapped models for $p$-torus data

We consider robust estimation of wrapped models to multivariate circular data that are points on the surface of a $p$-torus based on the weighted likelihood methodology.Robust model fitting is achieved by a set of weighted likelihood estimating equations, based on the computation of data dependent weights aimed to down-weight anomalous values, such as unexpected directions that do not share the main pattern of the bulk of the data. Weighted likelihood estimating equations with weights evaluated on the torus orobtained after unwrapping the data onto the Euclidean space are proposed and compared. Asymptotic properties and robustness features of the estimators under study have been studied, whereas their finite sample behavior has been investigated by Monte Carlo numerical experiment and real data examples.

stat.ME

A regularized MANOVA test for semicontinuous high-dimensional data

We propose a MANOVA test for semicontinuous data that is applicable also when the dimensionality exceeds the sample size. The test statistic is obtained as a likelihood ratio, where numerator and denominator are computed at the maxima of penalized likelihood functions under each hypothesis. Closed form solutions for the regularized estimators allow us to avoid computational overheads. We derive the null distribution using a permutation scheme. The power and level of the resulting test are evaluated in a simulation study. We illustrate the new methodology with two original data analyses, one regarding microRNA expression in human blastocyst cultures, and another regarding alien plant species invasion in the island of Socotra (Yemen).

stat.ME

Robust Penalized Estimators for High--Dimensional Generalized Linear Models

Robust estimators for generalized linear models (GLMs) are not easy to develop due to the nature of the distributions involved. Recently, there has been growing interest in robust estimation methods, particularly in contexts involving a potentially large number of explanatory variables. Transformed M-estimators (MT-estimators) provide a natural extension of M-estimation techniques to the GLM framework, offering robust methodologies. We propose a penalized variant of MT-estimators to address high-dimensional data scenarios. Under suitable assumptions, we demonstrate the consistency and asymptotic normality of this novel class of estimators. Our theoretical development focuses on redescending rho-functions and penalization functions that satisfy specific regularity conditions. We present an Iterative Re-Weighted Least Squares algorithm, together with a deterministic initialization procedure, which is crucial since the estimating equations may have multiple solutions. We evaluate the finite sample performance of this method for Poisson distribution and well known penalization functions through Monte Carlo simulations that consider various types of contamination, as well as an empirical application using a real dataset.

stat.ME

Temporally-Evolving Generalised Networks and their Reproducing Kernels

This paper considers generalised network, intended as networks where (a) the edges connecting the nodes are nonlinear, and (b) stochastic processes are continuously indexed over both vertices and edges. Such topological structures are normally represented through special classes of graphs, termed graphs with Euclidean edges. We build generalised networks in which topology changes over time instants. That is, vertices and edges can disappear at subsequent time instants and edges may change in shape and length. We consider both cases of linear or circular time. For the second case, the generalised network exhibits a periodic structure. Our findings allow to illustrate pros and cons of each setting. Generalised networks become semi-metric spaces whenever equipped with a proper semi-metric. Our approach allows to build proper semi-metrics for the temporally-evolving topological structures of the networks. Our final effort is then devoted to guiding the reader through appropriate choice of classes of functions that allow to build proper reproducing kernels when composed with the temporally-evolving semi-metrics topological structures.

cs.SI

Analytical and statistical properties of local depth functions motivated by clustering applications

Local general depth ($LGD$) functions are used for describing the local geometric features and mode(s) in multivariate distributions. In this paper, we undertake a rigorous systematic study of $LGD$ and establish several analytical and statistical properties. First, we show that, when the underlying probability distribution is absolutely continuous with density $f(\cdot)$, the scaled version of $LGD$ (referred to as $τ$-approximation) converges, uniformly and in $L^d(\mathbb{R}^p)$ to $f(\cdot)$ when $τ$ converges to zero. Second, we establish that, as the sample size diverges to infinity the centered and scaled sample $LGD$ converge in distribution to a centered Gaussian process uniformly in the space of bounded functions on $\mathcal{H}_G$, a class of functions yielding $LGD$. Third, using the sample version of the $τ$-approximation ($S τA$) and the gradient system analysis, we develop a new clustering algorithm. The validity of this algorithm requires several results concerning the uniform finite difference approximation of the gradient system associated with $S τA$. For this reason, we establish \emph{Bernstein}-type inequality for deviations between the centered and scaled sample $LGD$, which is also of independent interest. Finally, invoking the above results, we establish consistency of the clustering algorithm. Applications of the proposed methods to mode estimation and upper level set estimation are also provided. Finite sample performance of the methodology are evaluated using numerical experiments and data analysis.

math.ST

A Robust Seemingly Unrelated Regressions For Row-Wise And Cell-Wise Contamination

The Seemingly Unrelated Regressions (SUR) model is a wide used estimation procedure in econometrics, insurance and finance, where very often, the regression model contains more than one equation. Unknown parameters, regression coefficients and covariances among the errors terms, are estimated using algorithms based on Generalized Least Squares or Maximum Likelihood, and the method, as a whole, is very sensitive to outliers. To overcome this problem M-estimators and S-estimators are proposed in the literature together with fast algorithms. However, these procedures are only able to cope with row-wise outliers in the error terms, while their performance becomes very poor in the presence of cell-wise outliers and as the number of equations increases. A new robust approach is proposed which is able to perform well under both contamination types as well as it is fast to compute. Illustrations based on Monte Carlo simulations and a real data example are provided.

stat.ME

Robust Multivariate Estimation Based On Statistical Depth Filters

In the classical contamination models, such as the gross-error (Huber and Tukey contamination model or Case-wise Contamination), observations are considered as the units to be identified as outliers or not. This model is very useful when the number of considered variables is moderately small. Alqallaf et al. [2009] shows the limits of this approach for a larger number of variables and introduced the Independent contamination model (Cell-wise Contamination) where now the cells are the units to be identified as outliers or not. One approach to deal, at the same time, with both type of contamination is filter out the contaminated cells from the data set and then apply a robust procedure able to handle case-wise outliers and missing values. Here we develop a general framework to build filters in any dimension based on statistical data depth functions. We show that previous approaches, e.g. Agostinelli et al. [2015a] and Leung et al. [2017], are special cases. We illustrate our method by using the half-space depth.

math.ST

Robust Estimation for Multivariate Wrapped Models

A weighted likelihood technique for robust estimation of a multivariate Wrapped Normal distribution for data points scattered on a p-dimensional torus is proposed. The occurrence of outliers in the sample at hand can badly compromise inference for standard techniques such as maximum likelihood method. Therefore, there is the need to handle such model inadequacies in the fitting process by a robust technique and an effective down-weighting of observations not following the assumed model. Furthermore, the employ of a robust method could help in situations of hidden and unexpected substructures in the data. Here, it is suggested to build a set of data-dependent weights based on the Pearson residuals and solve the corresponding weighted likelihood estimating equations. In particular, robust estimation is carried out by using a Classification EM algorithm whose M-step is enhanced by the computation of weights based on current parameters' values. The finite sample behavior of the proposed method has been investigated by a Monte Carlo numerical studies and real data examples.

stat.ME

Robust Estimation under Linear Mixed Models: The Minimum Density Power Divergence Approach

Many real-life data sets can be analyzed using Linear Mixed Models (LMMs). Since these are ordinarily based on normality assumptions, under small deviations from the model the inference can be highly unstable when the associated parameters are estimated by classical methods. On the other hand, the density power divergence (DPD) family, which measures the discrepancy between two probability density functions, has been successfully used to build robust estimators with high stability associated with minimal loss in efficiency. Here, we develop the minimum DPD estimator (MDPDE) for independent but non identically distributed observations in LMMs. We prove the theoretical properties, including consistency and asymptotic normality. The influence function and sensitivity measures are studied to explore the robustness properties. As a data based choice of the MDPDE tuning parameter $\alpha$ is very important, we propose two candidates as "optimal" choices, where optimality is in the sense of choosing the strongest downweighting that is necessary for the particular data set. We conduct a simulation study comparing the proposed MDPDE, for different values of $\alpha$, with the S-estimators, M-estimators and the classical maximum likelihood estimator, considering different levels of contamination. Finally, we illustrate the performance of our proposal on a real-data example.

stat.ME

Torus Probabilistic Principal Component Analysis

Analyzing data in non-Euclidean spaces, such as bioinformatics, biology, and geology, where variables represent directions or angles, poses unique challenges. This type of data is known as circular data in univariate cases and can be termed spherical or toroidal in multivariate contexts. In this paper, we introduce a novel extension of Probabilistic Principal Component Analysis (PPCA) designed for toroidal (or torus) data, termed Torus Probabilistic PCA (TPPCA). We provide detailed algorithms for implementing TPPCA and demonstrate its applicability to torus data. To assess the efficacy of TPPCA, we perform comparative analyses using a simulation study and three real datasets. Our findings highlight the advantages and limitations of TPPCA in handling torus data. Furthermore, we propose statistical tests based on likelihood ratio statistics to determine the optimal number of components, enhancing the practical utility of TPPCA for real-world applications.

stat.AP