SearcharxivSearch

arXiv subjects

Paromita Banerjee

Publications and source records attributed to Paromita Banerjee.

4 recordsLinked to original sources

Robust K-means Clustering using the Density Power Divergence Measure

We introduce a robust clustering method, MK-means DPD, that estimates cluster centers and covariance matrices using density power divergence (DPD) measures combined with Mahalanobis distance, making it resistant to outliers and adaptable to heterogeneous, elliptical clusters, unlike the classical K-means algorithm. Since Mahalanobis distance-based K-means lacks a general convergence guarantee, we further introduce a convergent variant, Density-Consistent MK-means DPD (DC-MK-means DPD), which redefines the cluster assignment step in terms of a pointwise DPD loss. We prove a formal theorem establishing that the resulting algorithm converges in a finite number of steps. We also propose two new robust internal evaluation indices, a Median Davies-Bouldin Index and a Trimmed Calinski-Harabasz Index, to ensure that performance comparisons are not themselves distorted by outliers. The efficacy of the proposed methods is demonstrated on simulated data, showing superiority over existing methods, and on two real datasets: Iris data, to identify similar species, and COVID-19 case fatality rate and infection rate data for countries worldwide, examining the resulting clusters' geographic and socio-economic patterns.

stat.ME

A Bayesian Time-Varying SEIARD Model for State-Level COVID-19 Transmission and Mortality in the United States

We conduct a retrospective analysis of COVID-19 transmission dynamics across U.S. states using a modified population-based Susceptible-Exposed-Infectious-Asymptomatic-Recovered-Deceased (SEIARD) compartmental model. The proposed framework introduces time-varying transmission, reporting, and mortality rates to capture temporal variations in public behavior and policy interventions during the pandemic. In particular, the transmission rate is modeled as a function of population mobility (derived from Google Mobility Reports), with a residual time-decay term capturing the net effect of unobserved factors such as behavioral adaptation and control measures, while reporting is linked to nationwide testing strategies. We employ a Bayesian approach to integrate multiple data sources and quantify uncertainties in model parameters. The model explicitly distinguishes between symptomatic and asymptomatic infectious individuals and links the latent epidemic states to observable quantities, including reported cases and deaths, through a dynamic reporting function. This retrospective modeling framework provides insights into state-level epidemic trajectories and supports data-driven decision-making for optimal allocation of healthcare resources and evaluation of public health interventions during future pandemics. We further apply a clustering analysis to the posterior parameter estimates to identify groups of U.S. states exhibiting similar epidemiological characteristics, revealing substantial regional heterogeneity in transmission intensity, reproduction dynamics, and mortality burden.

stat.AP

Copula-Based Reconstruction and Clustering of Coccidioides Minimum Inhibitory Concentration Profiles

Coccidioidomycosis is a fungal lung infection endemic to parts of the Pacific Northwest and southwestern United States, Mexico, Central America, and South America. A 2017 study by Thompson et al. reported that many analyzed Coccidioides isolates had elevated minimum inhibitory concentration values for fluconazole but comparatively low minimum inhibitory concentration values for other triazole drugs. We constructed 172 synthetic joint minimum inhibitory concentration profiles from the published drug-specific marginal frequency tables. Values sampled from each marginal distribution were paired across drugs using a Gaussian copula, and potential cross-drug patterns were examined using k-means clustering, hierarchical clustering, and Gaussian mixture models. On the selected reconstruction, a forced four-cluster partition consistently identified an upper caspofungin minimum-inhibitory-concentration tail. Across the evaluated dependence scenarios, the gap statistic favored a single cluster in 96% to 100% of nested reconstructions, providing no evidence that the reconstructed multivariate data supported a broader multicluster structure. A separate simulation study examined recovery of an upper-severity-score category defined from prespecified quantiles of a composite minimum-inhibitory-concentration score. The study included 4,000 replications across 20 prevalence and measurement-noise conditions. The multiclass adjusted Rand index ranged from moderate to high depending on the method and prevalence, whereas binary partition agreement for the upper-severity-score category approached zero at 1% prevalence even though recall remained near 1.0. Clustering conclusions therefore depended on the assumed cross-drug dependence structure and on the realized number of observations in the target category.

stat.AP

An Improved Milstein Method for the Numerical Solution of Multidimensional Stochastic Differential Equations

Stochastic differential equations (SDEs) offer powerful and accessible mathematical models for capturing both deterministic and probabilistic aspects of dynamic behavior across a wide range of physical, financial, and social systems. However, analytical solutions for many SDEs are often unavailable, necessitating the use of numerical approximation methods. The rate of convergence of such numerical methods is of great importance, as it directly influences both computational efficiency and accuracy. This paper presents a proposed theorem, along with its proof, that facilitates the numerical evaluation of the strong (and weak) order of convergence of a numerical scheme for an SDE when the analytical solution is unavailable. Additionally, we address the challenge of numerically computing the multiple stochastic integrals required by the Milstein method to achieve improved convergence rates for multidimensional SDEs. In this context, two newly proposed numerical techniques for computing these multiple stochastic integrals are introduced and compared with existing approaches in terms of efficiency and effectiveness. The methodologies are further illustrated through simulation studies and applications to widely used financial models.

math.ST