SearcharxivSearch

arXiv subjects

Subhrajyoty Roy

Publications and source records attributed to Subhrajyoty Roy.

17 recordsLinked to original sources

Universally Optimal Robustness-Efficiency Tradeoffs for a General Class of Minimum Divergence Estimators

Balancing the efficiency of an estimator under ideal conditions against its robustness under contamination remains a central challenge in robust statistics. While minimum divergence methods offer a flexible alternative to traditional M-estimation, choosing the appropriate discrepancy measure has historically relied on heuristic or empirical justifications. This manuscript introduces a rigorous optimality criterion for this selection process. By investigating the comprehensive Generalized Alpha-Beta Divergence (GABD) family, we explicitly characterize the Pareto frontier dictating the lowest possible asymptotic variance for any strictly enforced asymptotic breakdown point. Our main theoretical results establish that the estimator achieving this mathematical optimum invariably falls within the extended $(\phi, \gamma)$-divergence class. Crucially, the derived optimal tuning parameter, $\phi^*$, given other parameters, depends solely on the desired breakdown threshold and is entirely invariant to both the assumed parametric model and the exact nature of the data contamination. Supported by comprehensive derivations of asymptotic normality, influence functions, and breakdown thresholds for both continuous and discrete settings, this work offers a unified, theoretical resolution to the long-standing problem of optimal divergence selection in robust inference.

math.ST

BOOOM: Loss-Function-Agnostic Black-Box Optimization over Orthonormal Manifolds for Machine Learning and Statistical Inference

Optimization over the Stiefel manifold $\mathrm{St}(p,d)$, the set of $p \times d$ column-orthonormal matrices, is fundamental in statistics, machine learning, and scientific computing, yet remains challenging in the presence of non-convex, non-smooth, or black-box objectives. Existing methods largely rely on either convex relaxations or gradient-based Riemannian optimization, limiting applicability in derivative-free and highly multimodal settings. We propose \textsc{BOOOM} (Black-box Optimization Over Orthonormal Manifolds), a general-purpose framework for loss-function-agnostic optimization on $\mathrm{St}(p,d)$. The key idea is a global Givens rotation-based parametrization that maps the manifold to an unconstrained Euclidean angle space while preserving feasibility exactly. Building on this representation, BOOOM employs a structured, parallelizable, derivative-free search based on Recursive Modified Pattern Search, enabling systematic exploration through plane-wise rotations without requiring gradient information and facilitating escape from poor local optima. We establish a unified theoretical framework showing equivalence between angle-space and manifold optimization, transfer of stationarity, and global convergence in probability under mild conditions. Empirical results across diverse problems, including heterogeneous quadratic optimization, low-rank and sparse matrix decomposition, independent component analysis, and orthogonal joint diagonalization, among other widely studied settings, demonstrate strong performance relative to state-of-the-art methods, particularly in non-smooth and highly multimodal regimes. We further illustrate its practical utility through a novel supervised PCA formulation applied to metabolomics data in colorectal cancer.

math.OC

Nonparametric regression of spatio-temporal data using infinite-dimensional covariates

In spatio-temporal analysis, we often record data at specific time intervals but with varying spatial locations between these timepoints. We propose a conditional model to analyze such spatio-temporal data that accommodates the dependencies alongside second-order stationary explanatory variables, which may be infinite-dimensional and accommodate spatio-temporal covariates. Because of the absence of a mixing-type dependence condition in this case, which is typically required by the existing studies, we consider a weaker polynomially decaying moment contraction (PMC) condition on the covariates. In this paper, we obtain nonparametric point estimates of the mean and covariate functions of such a regression model, which we then show to be statistically consistent. We also obtain a simultaneous confidence interval of the mean function using the central limit theorem for the proposed estimator. Such simultaneous inference tools can be used to test for certain specifications of the mean function. Some simulation studies and two real-data analyses have been illustrated to corroborate the findings.

stat.ME

Robust Rank Estimation for Noisy Matrices

Estimating the true rank of a noisy data matrix is a fundamental problem underlying techniques such as principal component analysis, matrix completion, etc. Existing rank estimation criteria, including information-based and cross-validation methods, are either highly sensitive to outliers or computationally demanding when combined with robust estimators. This paper proposes a new criterion, the Divergence Information Criterion for Matrix Rank (DICMR), that achieves both robustness and computational simplicity. Derived from the density power divergence framework, DICMR inherits the robustness properties while being computationally very simple. We provide asymptotic bounds on its overestimation and underestimation probabilities, and demonstrate first-order B-robustness of the criteria. Extensive simulations show that DICMR delivers accuracy comparable to the robustified cross-validation methods, but with far lower computational cost. We also showcase a real-data application to microarray imputation to further demonstrate its practical utility, outperforming several state-of-the-art algorithms.

stat.ME

Fast segmentation of watermarked texts from large language models through an epidemic change-point framework

With the growing use of large language models, concerns over content authenticity have spurred a variety of watermarking schemes. These schemes use secret keys to detect machine-generated text while remaining imperceptible to readers. Detection typically reduces to statistical hypothesis testing for the presence of watermarks, a topic that is now well studied. In contrast, the finer-grained task of localizing which segments of a text are watermarked is much less explored; existing approaches often lack scalability or guarantees robust to paraphrasing and post-editing. We bring a new perspective to this segmentation problem through the lens of epidemic change-points and, by exploiting this connection, propose WISER, a novel and computationally efficient watermark segmentation algorithm. We establish finite-sample error bounds and consistency for detecting multiple watermarked segments in a single text. Complementing these theoretical results, our extensive numerical experiments show that WISER outperforms state-of-the-art baseline methods, both in terms of computational speed as well as accuracy, on various benchmark datasets embedded with diverse watermarking schemes. Together, these theoretical and empirical results position WISER as an effective tool for watermark localization and illustrate how classical statistical ideas can yield theoretically valid and computationally efficient solutions to a modern problem of immediate importance.

stat.ML

Asymptotic breakdown point analysis of the minimum density power divergence estimator under independent non-homogeneous setups

The minimum density power divergence estimator (MDPDE) has gained significant attention in the literature of robust inference due to its strong robustness properties and high asymptotic efficiency; it is relatively easy to compute and can be interpreted as a generalization of the classical maximum likelihood estimator. It has been successfully applied in various setups, including the case of independent and non-homogeneous (INH) observations that cover both classification and regression-type problems with a fixed design. While the local robustness of this estimator has been theoretically validated through the bounded influence function, no general result is known about the global reliability or the breakdown behavior of this estimator under the INH setup, except for the specific case of location-type models. In this paper, we extend the notion of asymptotic breakdown point from the case of independent and identically distributed data to the INH setup and derive a theoretical lower bound for the asymptotic breakdown point of the MDPDE, under some easily verifiable assumptions. These results are further illustrated with applications to some fixed design regression models and corroborated through extensive simulation studies.

math.ST

Characterization of Generalized Alpha-Beta Divergence and Associated Entropy Measures

Minimum divergence estimators provide a natural framework for robust (parametric) statistical inference. Useful properties of several such divergence measures, including, the Hellinger distance, the power divergence, the density power divergence, the logarithmic density power divergence, etc., have been established in the literature; many of them lead to estimators with high statistical efficiency, sometimes even full asymptotic efficiency. The notable success of these divergences as tools of parametric inference motivates us to explore possible extensions of the alpha-beta divergence family, leading to a superfamily of divergence measures called the ``generalized alpha-beta (GAB) divergences''. This family contains all the aforementioned popular divergence measures as special cases, and additionally provides opportunities to discover new and novel classes of divergences that generate estimators having strong robustness properties without allowing a significant drop in statistical efficiency in various applications. In this paper, we provide the necessary and sufficient conditions for the validity of these generalized divergence measures that enable us to employ them for improved statistical inference. We also show various characterizing properties like duality, inversion, semi-continuity, etc., for the general class of GAB divergences. A discussion on the entropy measure derived from this general family and its properties are also presented along with the associated maximum entropy principle. The class of GAB divergences provide a delicate balance between local and global robustness, and this is illustrated by two examples of robust parameter estimation under the Geometric and the normal scale models.

math.ST

Asymptotic Breakdown Point Analysis for a General Class of Minimum Divergence Estimators

Robust inference based on the minimization of statistical divergences has proved to be a useful alternative to classical techniques based on maximum likelihood and related methods. Basu et al. (1998) introduced the density power divergence (DPD) family as a measure of discrepancy between two probability density functions and used this family for robust estimation of the parameter for independent and identically distributed data. Ghosh et al. (2017) proposed a more general class of divergence measures, namely the S-divergence family and discussed its usefulness in robust parametric estimation through several asymptotic properties and some numerical illustrations. In this paper, we develop the results concerning the asymptotic breakdown point for the minimum S-divergence estimators (in particular the minimum DPD estimator) under general model setups. The primary result of this paper provides lower bounds to the asymptotic breakdown point of these estimators which are independent of the dimension of the data, in turn corroborating their usefulness in robust inference under high dimensional data.

math.ST

Nonparametric quantile regression for spatio-temporal processes

In this paper, we develop a new and effective approach to nonparametric quantile regression that accommodates ultrahigh-dimensional data arising from spatio-temporal processes. This approach proves advantageous in staving off computational challenges that constitute known hindrances to existing nonparametric quantile regression methods when the number of predictors is much larger than the available sample size. We investigate conditions under which estimation is feasible and of good overall quality and obtain sharp approximations that we employ to devising statistical inference methodology. These include simultaneous confidence intervals and tests of hypotheses, whose asymptotics is borne by a non-trivial functional central limit theorem tailored to martingale differences. Additionally, we provide finite-sample results through various simulations which, accompanied by an illustrative application to real-worldesque data (on electricity demand), offer guarantees on the performance of the proposed methodology.

stat.ME

Trustworthy Dimensionality Reduction

Different unsupervised models for dimensionality reduction like PCA, LLE, Shannon's mapping, tSNE, UMAP, etc. work on different principles, hence, they are difficult to compare on the same ground. Although they are usually good for visualisation purposes, they can produce spurious patterns that are not present in the original data, losing its trustability (or credibility). On the other hand, information about some response variable (or knowledge of class labels) allows us to do supervised dimensionality reduction such as SIR, SAVE, etc. which work to reduce the data dimension without hampering its ability to explain the particular response at hand. Therefore, the reduced dataset cannot be used to further analyze its relationship with some other kind of responses, i.e., it loses its generalizability. To make a better dimensionality reduction algorithm with a better balance between these two, we shall formally describe the mathematical model used by dimensionality reduction algorithms and provide two indices to measure these intuitive concepts such as trustability and generalizability. Then, we propose a Localized Skeletonization and Dimensionality Reduction (LSDR) algorithm which approximately achieves optimality in both these indices to some extent. The proposed algorithm has been compared with state-of-the-art algorithms such as tSNE and UMAP and is found to be better overall in preserving global structure while retaining useful local information as well. We also propose some of the possible extensions of LSDR which could make this algorithm universally applicable for various types of data similar to tSNE and UMAP.

stat.ME

Robust and Efficient Estimation in Ordinal Response Models using the Density Power Divergence

In real life, we frequently come across data sets that involve some independent explanatory variable(s) generating a set of ordinal responses. These ordinal responses may correspond to an underlying continuous latent variable, which is linearly related to the covariate(s), and takes a particular (ordinal) label depending on whether this latent variable takes value in some suitable interval specified by a pair of (unknown) cut-offs. The most efficient way of estimating the unknown parameters (i.e., the regression coefficients and the cut-offs) is the method of maximum likelihood (ML). However, contamination in the data set either in the form of misspecification of ordinal responses, or the unboundedness of the covariate(s), might destabilize the likelihood function to a great extent where the ML based methodology might lead to completely unreliable inferences. In this paper, we explore a minimum distance estimation procedure based on the popular density power divergence (DPD) to yield robust parameter estimates for the ordinal response model. This paper highlights how the resulting estimator, namely the minimum DPD estimator (MDPDE), can be used as a practical robust alternative to the classical procedures based on the ML. We rigorously develop several theoretical properties of this estimator, and provide extensive simulations to substantiate the theory developed.

stat.ME

Robust Principal Component Analysis using Density Power Divergence

Principal component analysis (PCA) is a widely employed statistical tool used primarily for dimensionality reduction. However, it is known to be adversely affected by the presence of outlying observations in the sample, which is quite common. Robust PCA methods using M-estimators have theoretical benefits, but their robustness drop substantially for high dimensional data. On the other end of the spectrum, robust PCA algorithms solving principal component pursuit or similar optimization problems have high breakdown, but lack theoretical richness and demand high computational power compared to the M-estimators. We introduce a novel robust PCA estimator based on the minimum density power divergence estimator. This combines the theoretical strength of the M-estimators and the minimum divergence estimators with a high breakdown guarantee regardless of data dimension. We present a computationally efficient algorithm for this estimate. Our theoretical findings are supported by extensive simulations and comparisons with existing robust PCA methods. We also showcase the proposed algorithm's applicability on two benchmark datasets and a credit card transactions dataset for fraud detection.

stat.ME

rSVDdpd: A Robust Scalable Video Surveillance Background Modelling Algorithm

A basic algorithmic task in automated video surveillance is to separate background and foreground objects. Camera tampering, noisy videos, low frame rate, etc., pose difficulties in solving the problem. A general approach that classifies the tampered frames, and performs subsequent analysis on the remaining frames after discarding the tampered ones, results in loss of information. Several robust methods based on robust principal component analysis (PCA) have been introduced to solve this problem. To date, considerable effort has been expended to develop robust PCA via Principal Component Pursuit (PCP) methods with reduced computational cost and visually appealing foreground detection. However, the convex optimizations used in these algorithms do not scale well to real-world large datasets due to large matrix inversion steps. Also, an integral component of these foreground detection algorithms is singular value decomposition which is nonrobust. In this paper, we present a new video surveillance background modelling algorithm based on a new robust singular value decomposition technique rSVDdpd which takes care of both these issues. We also demonstrate the superiority of our proposed algorithm on a benchmark dataset and a new real-life video surveillance dataset in the presence of camera tampering. Software codes and additional illustrations are made available at the accompanying website rSVDdpd Homepage (https://subroy13.github.io/rsvddpd-home/)

stat.AP

A Generalized Epidemiological Model for COVID-19 with Dynamic and Asymptomatic Population

In this paper, we develop an extension of standard epidemiological models, suitable for COVID-19. This extension incorporates the transmission due to pre-symptomatic or asymptomatic carriers of the virus. Furthermore, this model also captures the spread of the disease due to the movement of people to/from different administrative boundaries within a country. The model describes the probabilistic rise in the number of confirmed cases due to the concomitant effects of (incipient) human transmission and multiple compartments. The associated parameters in the model can help architect the public health policy and operational management of the pandemic. For instance, this model demonstrates that increasing the testing for symptomatic patients does not have any major effect on the progression of the pandemic, but testing rate of the asymptomatic population has an extremely crucial role to play. The model is executed using the data obtained for the state of Chhattisgarh in the Republic of India. The model is shown to have significantly better predictive capability than the other epidemiological models. This model can be readily applied to any administrative boundary (state or country). Moreover, this model can be applied for any other epidemic as well.

q-bio.PE

Rough-Fuzzy CPD: A Gradual Change Point Detection Algorithm

Changepoint detection is the problem of finding abrupt or gradual changes in time series data when the distribution of the time series changes significantly. There are many sophisticated statistical algorithms for solving changepoint detection problem, although there is not much work devoted towards gradual changepoints as compared to abrupt ones. Here we present a new approach to solve changepoint detection problem using fuzzy rough set theory which is able to detect such gradual changepoints. An expression for the rough-fuzzy estimate of changepoints is derived along with its mathematical properties concerning fast computation. In a statistical hypothesis testing framework, asymptotic distribution of the proposed statistic on both single and multiple changepoints is derived under null hypothesis enabling multiple changepoint detection. Extensive simulation studies have been performed to investigate how simple crude statistical measures of disparity can be subjected to improve their efficiency in estimation of gradual changepoints. Also, the said rough-fuzzy estimate is robust to signal-to-noise ratio, high degree of fuzziness in true changepoints and also to hyper parameter values. Simulation studies reveal that the proposed method beats other fuzzy methods and also popular crisp methods like WBS, PELT and BOCD in detecting gradual changepoints. The applicability of the estimate is demonstrated using multiple real-life datasets including Covid-19. We have developed the python package "roufcp" for broader dissemination of the methods.

stat.ME

The Information Content of Taster's Valuation in Tea Auctions of India

Tea auctions across India occur as an ascending open auction, conducted online. Before the auction, a sample of the tea lot is sent to potential bidders and a group of tea tasters. The seller's reserve price is a confidential function of the tea taster's valuation, which also possibly acts as a signal to the bidders. In this paper, we work with the dataset from a single tea auction house, J Thomas, of tea dust category, on 49 weeks in the time span of 2018-2019, with the following objectives in mind: $\bullet$ Objective classification of the various categories of tea dust (25) into a more manageable, and robust classification of the tea dust, based on source and grades. $\bullet$ Predict which tea lots would be sold in the auction market, and a model for the final price conditioned on sale. $\bullet$ To study the distribution of price and ratio of the sold tea auction lots. $\bullet$ Make a detailed analysis of the information obtained from the tea taster's valuation and its impact on the final auction price. The model used has shown various promising results on cross-validation. The importance of valuation is firmly established through analysis of causal relationship between the valuation and the actual price. The authors hope that this study of the properties and the detailed analysis of the role played by the various factors, would be significant in the decision making process for the players of the auction game, pave the way to remove the manual interference in an attempt to automate the auction procedure, and improve tea quality in markets.

stat.AP

Onset detection: A new approach to QBH system

Query by Humming (QBH) is a system to provide a user with the song(s) which the user hums to the system. Current QBH method requires the extraction of onset and pitch information in order to track similarity with various versions of different songs. However, we here focus on detecting precise onsets only and use them to build a QBH system which is better than existing methods in terms of speed and memory and empirically in terms of accuracy. We also provide statistical analogy for onset detection functions and provide a measure of error in our algorithm.

stat.AP