SearcharxivSearch

arXiv subjects

Soham Bonnerjee

Publications and source records attributed to Soham Bonnerjee.

9 recordsLinked to original sources

Stability beyond Bounded Differences: Sharp Generalization Bounds under Finite $L_p$ Moments

While algorithmic stability is a central tool for understanding generalization of learning algorithms, existing high-probability guarantees typically rely on uniform boundedness or sub-Gaussian/sub-Weibull tail assumptions, which can be overly restrictive for modern settings with heavy-tailed or unbounded losses. We develop a stability-based framework that requires only a finite $L_p$ moment condition. Our first contribution is sharp concentration inequalities for functions of independent random variables under $L_p$ constraints, extending McDiarmid's bounded-differences techniques beyond the classical regime. Leveraging these results, we derive sharp high-probability generalization bounds across a range of learning paradigms, including empirical risk minimization, transductive regression, and meta-learning. These guarantees show that $L_p$ stability suffices for robust generalization even when boundedness fails, substantially weakening the standard assumptions in the stability literature.

stat.ML

Sharp asymptotic theory for Q-learning with LDTZ learning rate and its generalization

Despite the sustained popularity of Q-learning as a practical tool for policy determination, a majority of relevant theoretical literature deals with either constant ($\eta_{t}\equiv \eta$) or polynomially decaying ($\eta_{t} = \eta t^{-\alpha}$) learning schedules. However, it is well known that these choices suffer from either persistent bias or prohibitively slow convergence. In contrast, the recently proposed linear decay to zero (\texttt{LD2Z}: $\eta_{t,n}=\eta(1-t/n)$) schedule has shown appreciable empirical performance, but its theoretical and statistical properties remain largely unexplored, especially in the Q-learning setting. We address this gap in the literature by first considering a general class of power-law decay to zero (\texttt{PD2Z}-$\nu$: $\eta_{t,n}=\eta(1-t/n)^{\nu}$). Proceeding step-by-step, we present a sharp non-asymptotic error bound for Q-learning with \texttt{PD2Z}-$\nu$ schedule, which then is used to derive a central limit theory for a new \textit{tail} Polyak-Ruppert averaging estimator. Finally, we also provide a novel time-uniform Gaussian approximation (also known as \textit{strong invariance principle}) for the partial sum process of Q-learning iterates, which facilitates bootstrap-based inference. All our theoretical results are complemented by extensive numerical experiments. Beyond being new theoretical and statistical contributions to the Q-learning literature, our results definitively establish that \texttt{LD2Z} and in general \texttt{PD2Z}-$\nu$ achieve a best-of-both-worlds property: they inherit the rapid decay from initialization (characteristic of constant step-sizes) while retaining the asymptotic convergence guarantees (characteristic of polynomially decaying schedules). This dual advantage explains the empirical success of \texttt{LD2Z} while providing practical guidelines for inference through our results.

stat.ML

Fast localization of anomalous patches in spatial data under dependence

We propose a scalable, provably accurate method for localizing an unknown number of multiple axis-aligned anomalous patches in spatial data under a general class of spatial dependence. Motivated by the practical need to detect localized changes rather than completely segment large spatial grids, we first introduce both a naive and a significantly faster intelligent-sampling-based estimator for a single patch. We then extend this methodology to the highly challenging multiple-patch setting and propose a two-stage Spatial Patch Localization of Anomalies under DEpendence procedure (SPLADE). Under mild conditions on signal strength, separation from the boundary, inter-patch separation, and a uniform Gaussian approximation, we establish simultaneous consistency for the estimated number of patches and for each individual patch boundary. Extensive numerical results based on synthetic data scenarios demonstrate that the proposed method exhibits significant computational and accuracy gains over competing approaches, as well as robustness to moderate and severe spatial dependence. Finally, we demonstrate the real-world utility of the proposed method by applying it to frame-to-frame video surveillance data, where it accurately detects small, closely separated subjects, a task where existing methods are significantly slower and highly prone to spurious detections due to not accounting for spatial dependence. A second application on 3D fibrous media is deferred to the Appendix.

stat.ME

Fast segmentation of watermarked texts from large language models through an epidemic change-point framework

With the growing use of large language models, concerns over content authenticity have spurred a variety of watermarking schemes. These schemes use secret keys to detect machine-generated text while remaining imperceptible to readers. Detection typically reduces to statistical hypothesis testing for the presence of watermarks, a topic that is now well studied. In contrast, the finer-grained task of localizing which segments of a text are watermarked is much less explored; existing approaches often lack scalability or guarantees robust to paraphrasing and post-editing. We bring a new perspective to this segmentation problem through the lens of epidemic change-points and, by exploiting this connection, propose WISER, a novel and computationally efficient watermark segmentation algorithm. We establish finite-sample error bounds and consistency for detecting multiple watermarked segments in a single text. Complementing these theoretical results, our extensive numerical experiments show that WISER outperforms state-of-the-art baseline methods, both in terms of computational speed as well as accuracy, on various benchmark datasets embedded with diverse watermarking schemes. Together, these theoretical and empirical results position WISER as an effective tool for watermark localization and illustrate how classical statistical ideas can yield theoretically valid and computationally efficient solutions to a modern problem of immediate importance.

stat.ML

Sharp Gaussian approximations for Decentralized Federated Learning

Federated Learning has gained traction in privacy-sensitive collaborative environments, with local SGD emerging as a key optimization method in decentralized settings. While its convergence properties are well-studied, asymptotic statistical guarantees beyond convergence remain limited. In this paper, we present two generalized Gaussian approximation results for local SGD and explore their implications. First, we prove a Berry-Esseen theorem for the final local SGD iterates, enabling valid multiplier bootstrap procedures. Second, motivated by robustness considerations, we introduce two distinct time-uniform Gaussian approximations for the entire trajectory of local SGD. The time-uniform approximations support Gaussian bootstrap-based tests for detecting adversarial attacks. Extensive simulations are provided to support our theoretical results.

stat.ML

How Private is Your Attention? Bridging Privacy with In-Context Learning

In-context learning (ICL)-the ability of transformer-based models to perform new tasks from examples provided at inference time-has emerged as a hallmark of modern language models. While recent works have investigated the mechanisms underlying ICL, its feasibility under formal privacy constraints remains largely unexplored. In this paper, we propose a differentially private pretraining algorithm for linear attention heads and present the first theoretical analysis of the privacy-accuracy trade-off for ICL in linear regression. Our results characterize the fundamental tension between optimization and privacy-induced noise, formally capturing behaviors observed in private training via iterative methods. Additionally, we show that our method is robust to adversarial perturbations of training prompts, unlike standard ridge regression. All theoretical findings are supported by extensive simulations across diverse settings.

stat.ML

Gaussian Approximation For Non-stationary Time Series with Optimal Rate and Explicit Construction

Statistical inference for time series such as curve estimation for time-varying models or testing for existence of change-point have garnered significant attention. However, these works are generally restricted to the assumption of independence and/or stationarity at its best. The main obstacle is that the existing Gaussian approximation results for non-stationary processes only provide an existential proof and thus they are difficult to apply. In this paper, we provide two clear paths to construct such a Gaussian approximation for non-stationary series. While the first one is theoretically more natural, the second one is practically implementable. Our Gaussian approximation results are applicable for a very large class of non-stationary time series, obtain optimal rates and yet have good applicability. Building on such approximations, we also show theoretical results for change-point detection and simultaneous inference in presence of non-stationary errors. Finally we substantiate our theoretical results with simulation studies and real data analysis.

math.ST

A Generalized Epidemiological Model for COVID-19 with Dynamic and Asymptomatic Population

In this paper, we develop an extension of standard epidemiological models, suitable for COVID-19. This extension incorporates the transmission due to pre-symptomatic or asymptomatic carriers of the virus. Furthermore, this model also captures the spread of the disease due to the movement of people to/from different administrative boundaries within a country. The model describes the probabilistic rise in the number of confirmed cases due to the concomitant effects of (incipient) human transmission and multiple compartments. The associated parameters in the model can help architect the public health policy and operational management of the pandemic. For instance, this model demonstrates that increasing the testing for symptomatic patients does not have any major effect on the progression of the pandemic, but testing rate of the asymptomatic population has an extremely crucial role to play. The model is executed using the data obtained for the state of Chhattisgarh in the Republic of India. The model is shown to have significantly better predictive capability than the other epidemiological models. This model can be readily applied to any administrative boundary (state or country). Moreover, this model can be applied for any other epidemic as well.

q-bio.PE

Onset detection: A new approach to QBH system

Query by Humming (QBH) is a system to provide a user with the song(s) which the user hums to the system. Current QBH method requires the extraction of onset and pitch information in order to track similarity with various versions of different songs. However, we here focus on detecting precise onsets only and use them to build a QBH system which is better than existing methods in terms of speed and memory and empirically in terms of accuracy. We also provide statistical analogy for onset detection functions and provide a measure of error in our algorithm.

stat.AP