SearcharxivSearch

arXiv subjects

Ayanendranath Basu

Publications and source records attributed to Ayanendranath Basu.

At least 19 recordsLinked to original sources

A Composite Divergence Approach to Robust Multivariate Estimation under Cellwise and Casewise Contamination

Composite likelihood (CL) methods provide a computationally efficient alternative to full likelihood inference for complex multivariate models by replacing the joint likelihood with a product of lower-dimensional marginal or conditional components. Like the MLE, however, the maximum CL estimator (MCLE) is highly sensitive to data contamination. On the other hand, robust divergence-based procedures such as the minimum density power divergence (DPD) estimator require the full joint density and so scale poorly to complex multivariate models. We introduce the composite DPD (CDPD), a genuine statistical divergence built entirely from the low-dimensional component densities defining a CL, combining the computational scalability of CL with the robustness of the DPD. The resulting minimum CDPD estimator (MCDPDE) robustifies the MCLE without requiring integration over the full multivariate sample space. We establish consistency, asymptotic normality, and the influence function of the MCDPDE under regularity conditions on the component models alone, without requiring correct specification of the full joint distribution. We show that it is qualitatively robust for every positive value of its tuning parameter, unlike the MCLE recovered as the limit. Because its components can be chosen at the pairwise or cell level, the framework guards simultaneously against casewise and cellwise contamination. Operating directly on component densities rather than elliptical distance structures, it extends robust inference beyond the elliptical models to which most existing cellwise-robust procedures are confined. We develop computational algorithms implemented in the accompanying R package mvdpd. Simulation studies and real-data applications show that the MCDPDE achieves substantial robustness gains over the MCLE while retaining competitive efficiency under the assumed model.

math.ST

Universally Optimal Robustness-Efficiency Tradeoffs for a General Class of Minimum Divergence Estimators

Balancing the efficiency of an estimator under ideal conditions against its robustness under contamination remains a central challenge in robust statistics. While minimum divergence methods offer a flexible alternative to traditional M-estimation, choosing the appropriate discrepancy measure has historically relied on heuristic or empirical justifications. This manuscript introduces a rigorous optimality criterion for this selection process. By investigating the comprehensive Generalized Alpha-Beta Divergence (GABD) family, we explicitly characterize the Pareto frontier dictating the lowest possible asymptotic variance for any strictly enforced asymptotic breakdown point. Our main theoretical results establish that the estimator achieving this mathematical optimum invariably falls within the extended $(\phi, \gamma)$-divergence class. Crucially, the derived optimal tuning parameter, $\phi^*$, given other parameters, depends solely on the desired breakdown threshold and is entirely invariant to both the assumed parametric model and the exact nature of the data contamination. Supported by comprehensive derivations of asymptotic normality, influence functions, and breakdown thresholds for both continuous and discrete settings, this work offers a unified, theoretical resolution to the long-standing problem of optimal divergence selection in robust inference.

math.ST

Semiparametric Robust Estimation of Population Location

Real-world measurements often comprise a dominant signal contaminated by a noisy background. Robustly estimating the dominant signal in practice has been a fundamental statistical problem. Classically, mixture models have been used to cluster the heterogeneous population into homogeneous components. Modeling such data with fully parametric models risks bias under misspecification, while fully nonparametric approaches can dissipate power and computational resources. We propose a middle path: a semiparametric method that models only the dominant component parametrically and leaves the background completely nonparametric, yet remains computationally scalable and statistically robust. So instead of outlier downweighting, traditionally done in robust statistics literature, we maximize the observed likelihood such that the noisy background is absorbed by the nonparametric component. Computationally, we propose a new approximate FFT-accelerated likelihood maximization algorithm. Empirically, this FFT plug-in achieves order-of-magnitude speedups over vanilla weighted EM while preserving statistical accuracy and large sample properties.

stat.CO

Robust Rank Estimation for Noisy Matrices

Estimating the true rank of a noisy data matrix is a fundamental problem underlying techniques such as principal component analysis, matrix completion, etc. Existing rank estimation criteria, including information-based and cross-validation methods, are either highly sensitive to outliers or computationally demanding when combined with robust estimators. This paper proposes a new criterion, the Divergence Information Criterion for Matrix Rank (DICMR), that achieves both robustness and computational simplicity. Derived from the density power divergence framework, DICMR inherits the robustness properties while being computationally very simple. We provide asymptotic bounds on its overestimation and underestimation probabilities, and demonstrate first-order B-robustness of the criteria. Extensive simulations show that DICMR delivers accuracy comparable to the robustified cross-validation methods, but with far lower computational cost. We also showcase a real-data application to microarray imputation to further demonstrate its practical utility, outperforming several state-of-the-art algorithms.

stat.ME

Asymptotic breakdown point analysis of the minimum density power divergence estimator under independent non-homogeneous setups

The minimum density power divergence estimator (MDPDE) has gained significant attention in the literature of robust inference due to its strong robustness properties and high asymptotic efficiency; it is relatively easy to compute and can be interpreted as a generalization of the classical maximum likelihood estimator. It has been successfully applied in various setups, including the case of independent and non-homogeneous (INH) observations that cover both classification and regression-type problems with a fixed design. While the local robustness of this estimator has been theoretically validated through the bounded influence function, no general result is known about the global reliability or the breakdown behavior of this estimator under the INH setup, except for the specific case of location-type models. In this paper, we extend the notion of asymptotic breakdown point from the case of independent and identically distributed data to the INH setup and derive a theoretical lower bound for the asymptotic breakdown point of the MDPDE, under some easily verifiable assumptions. These results are further illustrated with applications to some fixed design regression models and corroborated through extensive simulation studies.

math.ST

A Weighted Likelihood Approach Based on Statistical Data Depths

We propose a general approach to construct weighted likelihood estimating equations with the aim of obtaining robust parameter estimates. We modify the standard likelihood equations by incorporating a weight that reflects the statistical depth of each data point relative to the model, as opposed to the sample. An observation is considered regular when the corresponding difference of these two depths is close to zero. When this difference is large the observation score contribution is downweighted. We study the asymptotic properties of the proposed estimator, including consistency and asymptotic normality, for a broad class of weight functions. In particular, we establish asymptotic normality under the standard regularity conditions typically assumed for the maximum likelihood estimator (MLE). Our weighted likelihood estimator achieves the same asymptotic efficiency as the MLE in the absence of contamination, while maintaining a high degree of robustness in contaminated settings. In stark contrast to the traditional minimum divergence/disparity estimators, our results hold even if the dimension of the data diverges with the sample size, without requiring additional assumptions on the existence or smoothness of the underlying densities. We also derive the finite sample breakdown point of our estimator for both location and scatter matrix in the elliptically symmetric model. Detailed results and examples are presented for robust parameter estimation in the multivariate normal model. Robustness is further illustrated using two real data sets and a Monte Carlo simulation study.

math.ST

Characterization of Generalized Alpha-Beta Divergence and Associated Entropy Measures

Minimum divergence estimators provide a natural framework for robust (parametric) statistical inference. Useful properties of several such divergence measures, including, the Hellinger distance, the power divergence, the density power divergence, the logarithmic density power divergence, etc., have been established in the literature; many of them lead to estimators with high statistical efficiency, sometimes even full asymptotic efficiency. The notable success of these divergences as tools of parametric inference motivates us to explore possible extensions of the alpha-beta divergence family, leading to a superfamily of divergence measures called the ``generalized alpha-beta (GAB) divergences''. This family contains all the aforementioned popular divergence measures as special cases, and additionally provides opportunities to discover new and novel classes of divergences that generate estimators having strong robustness properties without allowing a significant drop in statistical efficiency in various applications. In this paper, we provide the necessary and sufficient conditions for the validity of these generalized divergence measures that enable us to employ them for improved statistical inference. We also show various characterizing properties like duality, inversion, semi-continuity, etc., for the general class of GAB divergences. A discussion on the entropy measure derived from this general family and its properties are also presented along with the associated maximum entropy principle. The class of GAB divergences provide a delicate balance between local and global robustness, and this is illustrated by two examples of robust parameter estimation under the Geometric and the normal scale models.

math.ST

A Componentwise Estimation Procedure for Multivariate Location and Scatter: Robustness, Efficiency and Scalability

Covariance matrix estimation is an important problem in multivariate data analysis, both from theoretical as well as applied points of view. Many simple and popular covariance matrix estimators are known to be severely affected by model misspecification and the presence of outliers in the data; on the other hand robust estimators with reasonably high efficiency are often computationally challenging for modern large and complex datasets. In this work, we propose a new, simple, robust and highly efficient method for estimation of the location vector and the scatter matrix for elliptically symmetric distributions. The proposed estimation procedure is designed in the spirit of the minimum density power divergence (DPD) estimation approach with appropriate modifications which makes our proposal (sequential minimum DPD estimation) computationally very economical and scalable to large as well as higher dimensional datasets. Consistency and asymptotic normality of the proposed sequential estimators of the multivariate location and scatter are established along with asymptotic positive definiteness of the estimated scatter matrix. Robustness of our estimators are studied by means of influence functions. All theoretical results are illustrated further under multivariate normality. A large-scale simulation study is presented to assess finite sample performances and scalability of our method in comparison to the usual maximum likelihood estimator (MLE), the ordinary minimum DPD estimator (MDPDE) and other popular non-parametric methods. The applicability of our method is further illustrated with a real dataset on credit card transactions.

stat.ME

Robust inference for linear regression models with possibly skewed error distribution

Traditional methods for linear regression generally assume that the underlying error distribution, equivalently the distribution of the responses, is normal. Yet, sometimes real life response data may exhibit a skewed pattern, and assuming normality would not give reliable results in such cases. This is often observed in cases of some biomedical, behavioral, socio-economic and other variables. In this paper, we propose to use the class of skew normal (SN) distributions, which also includes the ordinary normal distribution as its special case, as the model for the errors in a linear regression setup and perform subsequent statistical inference using the popular and robust minimum density power divergence approach to get stable insights in the presence of possible data contamination (e.g., outliers). We provide the asymptotic distribution of the proposed estimator of the regression parameters and also propose robust Wald-type tests of significance for these parameters. We provide an influence function analysis of these estimators and test statistics, and also provide level and power influence functions. Numerical verification including simulation studies and real data analysis is provided to substantiate the theory developed.

stat.ME

Robust and Efficient Estimation in Ordinal Response Models using the Density Power Divergence

In real life, we frequently come across data sets that involve some independent explanatory variable(s) generating a set of ordinal responses. These ordinal responses may correspond to an underlying continuous latent variable, which is linearly related to the covariate(s), and takes a particular (ordinal) label depending on whether this latent variable takes value in some suitable interval specified by a pair of (unknown) cut-offs. The most efficient way of estimating the unknown parameters (i.e., the regression coefficients and the cut-offs) is the method of maximum likelihood (ML). However, contamination in the data set either in the form of misspecification of ordinal responses, or the unboundedness of the covariate(s), might destabilize the likelihood function to a great extent where the ML based methodology might lead to completely unreliable inferences. In this paper, we explore a minimum distance estimation procedure based on the popular density power divergence (DPD) to yield robust parameter estimates for the ordinal response model. This paper highlights how the resulting estimator, namely the minimum DPD estimator (MDPDE), can be used as a practical robust alternative to the classical procedures based on the ML. We rigorously develop several theoretical properties of this estimator, and provide extensive simulations to substantiate the theory developed.

stat.ME

Robust Clustering with Normal Mixture Models: A Pseudo $β$-Likelihood Approach

As in other estimation scenarios, likelihood based estimation in the normal mixture set-up is highly non-robust against model misspecification and presence of outliers (apart from being an ill-posed optimization problem). A robust alternative to the ordinary likelihood approach for this estimation problem is proposed which performs simultaneous estimation and data clustering and leads to subsequent anomaly detection. To invoke robustness, the methodology based on the minimization of the density power divergence (or alternatively, the maximization of the $β$-likelihood) is utilized under suitable constraints. An iteratively reweighted least squares approach has been followed in order to compute the proposed estimators for the component means (or equivalently cluster centers) and component dispersion matrices which leads to simultaneous data clustering. Some exploratory techniques are also suggested for anomaly detection, a problem of great importance in the domain of statistics and machine learning. The proposed method is validated with simulation studies under different set-ups; it performs competitively or better compared to the popular existing methods like K-medoids, TCLUST, trimmed K-means and MCLUST, especially when the mixture components (i.e., the clusters) share regions with significant overlap or outlying clusters exist with small but non-negligible weights (particularly in higher dimensions). Two real datasets are also used to illustrate the performance of the newly proposed method in comparison with others along with an application in image processing. The proposed method detects the clusters with lower misclassification rates and successfully points out the outlying (anomalous) observations from these datasets.

stat.ME

Robust Principal Component Analysis using Density Power Divergence

Principal component analysis (PCA) is a widely employed statistical tool used primarily for dimensionality reduction. However, it is known to be adversely affected by the presence of outlying observations in the sample, which is quite common. Robust PCA methods using M-estimators have theoretical benefits, but their robustness drop substantially for high dimensional data. On the other end of the spectrum, robust PCA algorithms solving principal component pursuit or similar optimization problems have high breakdown, but lack theoretical richness and demand high computational power compared to the M-estimators. We introduce a novel robust PCA estimator based on the minimum density power divergence estimator. This combines the theoretical strength of the M-estimators and the minimum divergence estimators with a high breakdown guarantee regardless of data dimension. We present a computationally efficient algorithm for this estimate. Our theoretical findings are supported by extensive simulations and comparisons with existing robust PCA methods. We also showcase the proposed algorithm's applicability on two benchmark datasets and a credit card transactions dataset for fraud detection.

stat.ME

rSVDdpd: A Robust Scalable Video Surveillance Background Modelling Algorithm

A basic algorithmic task in automated video surveillance is to separate background and foreground objects. Camera tampering, noisy videos, low frame rate, etc., pose difficulties in solving the problem. A general approach that classifies the tampered frames, and performs subsequent analysis on the remaining frames after discarding the tampered ones, results in loss of information. Several robust methods based on robust principal component analysis (PCA) have been introduced to solve this problem. To date, considerable effort has been expended to develop robust PCA via Principal Component Pursuit (PCP) methods with reduced computational cost and visually appealing foreground detection. However, the convex optimizations used in these algorithms do not scale well to real-world large datasets due to large matrix inversion steps. Also, an integral component of these foreground detection algorithms is singular value decomposition which is nonrobust. In this paper, we present a new video surveillance background modelling algorithm based on a new robust singular value decomposition technique rSVDdpd which takes care of both these issues. We also demonstrate the superiority of our proposed algorithm on a benchmark dataset and a new real-life video surveillance dataset in the presence of camera tampering. Software codes and additional illustrations are made available at the accompanying website rSVDdpd Homepage (https://subroy13.github.io/rsvddpd-home/)

stat.AP

Asymptotic Breakdown Point Analysis for a General Class of Minimum Divergence Estimators

Robust inference based on the minimization of statistical divergences has proved to be a useful alternative to classical techniques based on maximum likelihood and related methods. Basu et al. (1998) introduced the density power divergence (DPD) family as a measure of discrepancy between two probability density functions and used this family for robust estimation of the parameter for independent and identically distributed data. Ghosh et al. (2017) proposed a more general class of divergence measures, namely the S-divergence family and discussed its usefulness in robust parametric estimation through several asymptotic properties and some numerical illustrations. In this paper, we develop the results concerning the asymptotic breakdown point for the minimum S-divergence estimators (in particular the minimum DPD estimator) under general model setups. The primary result of this paper provides lower bounds to the asymptotic breakdown point of these estimators which are independent of the dimension of the data, in turn corroborating their usefulness in robust inference under high dimensional data.

math.ST

Robust adaptive Lasso in high-dimensional logistic regression

Penalized logistic regression is extremely useful for binary classification with large number of covariates (higher than the sample size), having several real life applications, including genomic disease classification. However, the existing methods based on the likelihood loss function are sensitive to data contamination and other noise and, hence, robust methods are needed for stable and more accurate inference. In this paper, we propose a family of robust estimators for sparse logistic models utilizing the popular density power divergence based loss function and the general adaptively weighted LASSO penalties. We study the local robustness of the proposed estimators through its influence function and also derive its oracle properties and asymptotic distribution. With extensive empirical illustrations, we demonstrate the significantly improved performance of our proposed estimators over the existing ones with particular gain in robustness. Our proposal is finally applied to analyse four different real datasets for cancer classification, obtaining robust and accurate models, that simultaneously performs gene selection and patient classification.

stat.ME

Existence and Consistency of the Maximum Pseudo \b{eta}-Likelihood Estimators for Multivariate Normal Mixture Models

Robust estimation under multivariate normal (MVN) mixture model is always a computational challenge. A recently proposed maximum pseudo \b{eta}-likelihood estimator aims to estimate the unknown parameters of a MVN mixture model in the spirit of minimum density power divergence (DPD) methodology but with a relatively simpler and tractable computational algorithm even for larger dimensions. In this letter, we will rigorously derive the existence and weak consistency of the maximum pseudo \b{eta}-likelihood estimator in case of MVN mixture models under a reasonable set of assumptions.

math.ST

Characterizing Logarithmic Bregman Functions

Minimum divergence procedures based on the density power divergence and the logarithmic density power divergence have been extremely popular and successful in generating inference procedures which combine a high degree of model efficiency with strong outlier stability. Such procedures are always preferable in practical situations over procedures which achieve their robustness at a major cost of efficiency or are highly efficient but have poor robustness properties. The density power divergence (DPD) family of Basu et al.(1998) and the logarithmic density power divergence (LDPD) family of Jones et al.(2001) provide flexible classes of divergences where the adjustment between efficiency and robustness is controlled by a single, real, non-negative parameter. The usefulness of these two families of divergences in statistical inference makes it meaningful to search for other related families of divergences in the same spirit. The DPD family is a member of the class of Bregman divergences, and the LDPD family is obtained by log transformations of the different segments of the divergences within the DPD family. Both the DPD and LDPD families lead to the Kullback-Leibler divergence in the limiting case as the tuning parameter $α\rightarrow 0$. In this paper we study this relation in detail, and demonstrate that such log transformations can only be meaningful in the context of the DPD (or the convex generating function of the DPD) within the general fold of Bregman divergences, giving us a limit to the extent to which the search for useful divergences could be successful.

math.ST

Characterizing the Functional Density Power Divergence Class

Divergence measures have a long association with statistical inference, machine learning and information theory. The density power divergence and related measures have produced many useful (and popular) statistical procedures, which provide a good balance between model efficiency on one hand and outlier stability or robustness on the other. The logarithmic density power divergence, a particular logarithmic transform of the density power divergence, has also been very successful in producing efficient and stable inference procedures; in addition it has also led to significant demonstrated applications in information theory. The success of the minimum divergence procedures based on the density power divergence and the logarithmic density power divergence (which also go by the names $β$-divergence and $γ$-divergence, respectively) make it imperative and meaningful to look for other, similar divergences which may be obtained as transforms of the density power divergence in the same spirit. With this motivation we search for such transforms of the density power divergence, referred to herein as the functional density power divergence class. The present article characterizes this functional density power divergence class, and thus identifies the available divergence measures within this construct that may be explored further for possible applications in statistical inference, machine learning and information theory.

math.ST