SearcharxivSearch

arXiv subjects

Sayar Karmakar

Publications and source records attributed to Sayar Karmakar.

At least 19 recordsLinked to original sources

Fast segmentation of watermarked texts from large language models through an epidemic change-point framework

With the growing use of large language models, concerns over content authenticity have spurred a variety of watermarking schemes. These schemes use secret keys to detect machine-generated text while remaining imperceptible to readers. Detection typically reduces to statistical hypothesis testing for the presence of watermarks, a topic that is now well studied. In contrast, the finer-grained task of localizing which segments of a text are watermarked is much less explored; existing approaches often lack scalability or guarantees robust to paraphrasing and post-editing. We bring a new perspective to this segmentation problem through the lens of epidemic change-points and, by exploiting this connection, propose WISER, a novel and computationally efficient watermark segmentation algorithm. We establish finite-sample error bounds and consistency for detecting multiple watermarked segments in a single text. Complementing these theoretical results, our extensive numerical experiments show that WISER outperforms state-of-the-art baseline methods, both in terms of computational speed as well as accuracy, on various benchmark datasets embedded with diverse watermarking schemes. Together, these theoretical and empirical results position WISER as an effective tool for watermark localization and illustrate how classical statistical ideas can yield theoretically valid and computationally efficient solutions to a modern problem of immediate importance.

stat.ML

Sharp Gaussian approximations for Decentralized Federated Learning

Federated Learning has gained traction in privacy-sensitive collaborative environments, with local SGD emerging as a key optimization method in decentralized settings. While its convergence properties are well-studied, asymptotic statistical guarantees beyond convergence remain limited. In this paper, we present two generalized Gaussian approximation results for local SGD and explore their implications. First, we prove a Berry-Esseen theorem for the final local SGD iterates, enabling valid multiplier bootstrap procedures. Second, motivated by robustness considerations, we introduce two distinct time-uniform Gaussian approximations for the entire trajectory of local SGD. The time-uniform approximations support Gaussian bootstrap-based tests for detecting adversarial attacks. Extensive simulations are provided to support our theoretical results.

stat.ML

Joint Estimation in Potts Model

In this paper, we study estimation of parameters in a two-parameter Potts model with $q$ colors and coupling matrix $A_N$. We characterize concrete sufficient conditions for existence of the pseudo-likelihood estimator of the Potts model, in terms of the local magnetic fields, and give sufficient conditions for the validity of the above characterization. We then provide sufficient criteria for estimation of both parameters at the optimal rate $\sqrt{N}$. In particular, if $A_N$ is the scaled adjacency matrix of a graph $G_N$, then we show that joint estimation is possible if either $G_N$ has bounded degree or is irregular. In contrast, we give an example of a graph sequence $G_N$ which is approximately regular and dense, where no consistent estimator exists. We also show that one-parameter estimation at the optimal rate $\sqrt{N}$ holds under much milder conditions when the other parameter is known. Along the way, we develop a concentration result for mean-field Potts models using the framework of nonlinear large deviations. Compared to the Ising case, our results for the Potts case require a novel analysis across multiple colors.

math.ST

Nonparametric regression of spatio-temporal data using infinite-dimensional covariates

In spatio-temporal analysis, we often record data at specific time intervals but with varying spatial locations between these timepoints. We propose a conditional model to analyze such spatio-temporal data that accommodates the dependencies alongside second-order stationary explanatory variables, which may be infinite-dimensional and accommodate spatio-temporal covariates. Because of the absence of a mixing-type dependence condition in this case, which is typically required by the existing studies, we consider a weaker polynomially decaying moment contraction (PMC) condition on the covariates. In this paper, we obtain nonparametric point estimates of the mean and covariate functions of such a regression model, which we then show to be statistically consistent. We also obtain a simultaneous confidence interval of the mean function using the central limit theorem for the proposed estimator. Such simultaneous inference tools can be used to test for certain specifications of the mean function. Some simulation studies and two real-data analyses have been illustrated to corroborate the findings.

stat.ME

Fast localization of anomalous patches in spatial data under dependence

We propose a scalable, provably accurate method for localizing an unknown number of multiple axis-aligned anomalous patches in spatial data under a general class of spatial dependence. Motivated by the practical need to detect localized changes rather than completely segment large spatial grids, we first introduce both a naive and a significantly faster intelligent-sampling-based estimator for a single patch. We then extend this methodology to the highly challenging multiple-patch setting and propose a two-stage Spatial Patch Localization of Anomalies under DEpendence procedure (SPLADE). Under mild conditions on signal strength, separation from the boundary, inter-patch separation, and a uniform Gaussian approximation, we establish simultaneous consistency for the estimated number of patches and for each individual patch boundary. Extensive numerical results based on synthetic data scenarios demonstrate that the proposed method exhibits significant computational and accuracy gains over competing approaches, as well as robustness to moderate and severe spatial dependence. Finally, we demonstrate the real-world utility of the proposed method by applying it to frame-to-frame video surveillance data, where it accurately detects small, closely separated subjects, a task where existing methods are significantly slower and highly prone to spurious detections due to not accounting for spatial dependence. A second application on 3D fibrous media is deferred to the Appendix.

stat.ME

Generalized percolation games on the $2$-dimensional square lattice, and ergodicity of associated probabilistic cellular automata

Each vertex of the infinite $2$-dimensional square lattice graph is assigned, independently, a label that reads trap with probability $p$, target with probability $q$, and open with probability $(1-p-q)$, and each edge is assigned, independently, a label that reads trap with probability $r$ and open with probability $(1-r)$. A percolation game is played on this random board, wherein two players take turns to make moves, where a move involves relocating the token from where it is currently located, say $(x,y) \in \mathbb{Z}^{2}$, to one of $(x+1,y)$ and $(x,y+1)$. A player wins if she is able to move the token to a vertex labeled a target, or force her opponent to either move the token to a vertex labeled a trap or along an edge labeled a trap. We seek to find a regime, in terms of $p$, $q$ and $r$, in which the probability of this game resulting in a draw equals $0$. We consider special cases of this game, such as when each edge is assigned, independently, a label that reads trap with probability $r$, target with probability $s$, and open with probability $(1-r-s)$, but the vertices are left unlabeled. Various regimes of values of $r$ and $s$ are explored in which the probability of draw is guaranteed to be $0$. We show that the probability of draw in each such game equals $0$ if and only if a certain probabilistic cellular automaton (PCA) is ergodic, following which we implement the technique of weight functions to investigate the regimes in which said PCA is ergodic.

math.PR

The Game of Graph Nim on graphs with four edges

This work is concerned with the study of the Game of Graph Nim -- a class of two-player combinatorial games -- on graphs with $4$ edges. To each edge of such a graph is assigned a positive-integer-valued edge-weight, and during each round of the game, the player whose turn it is to make a move selects a vertex, and removes a non-negative integer edge-weight from each of the edges incident on that vertex, such that the remaining edge-weight on each of these edges is a non-negative integer, and the total edge-weight removed during a round is strictly positive. The game continues for as long as the sum of the edge-weights remaining on all edges of the graph is strictly positive, and the player who plays the last round wins. An initial configuration of edge-weights is considered winning if the player who plays the first round wins the game, whereas it is defined as losing if the player who plays the second round wins. In this paper, we characterize, almost entirely, all winning and losing configurations for this game on all graphs with precisely $4$ edges each. Only one such graph defies our attempt to fully characterize the winning and losing configurations of edge-weights on its edges -- we are still able to provide a significant set of partial results pertaining to this graph.

math.CO

Percolation games on rooted, edge-weighted random trees

Consider a rooted Galton-Watson tree $T$, to each of whose edges we assign, independently, a weight that equals $+1$ with probability $p_{1}$, $0$ with probability $p_{0}$ and $-1$ with probability $p_{-1}=1-p_{1}-p_{0}$. We play a game on a realization of this tree, involving two players and a token that is allowed to be moved from where it is currently located, say a vertex $u$ of $T$, to any child $v$ of $u$. The players begin with initial capitals that amount to $i$ and $j$ units respectively, and a player wins if either she is the first to amass a capital worth $κ$ units, where $κ\in \mathbb{N}$ is prespecified, or she is able to move the token to a leaf vertex, from where her opponent cannot move it any farther, or her opponent's capital is the first to dwindle to $0$. This paper is concerned with analyzing the probabilities of the three possible outcomes such a game may culminate in, as well as with finding conditions under which the expected duration of the game is finite. Of particular interest to us is the exploration of criteria that guarantee the probability of draw in such a game to be $0$. The theory we develop is further supported by observations obtained via computer simulations, providing a deeper insight into how the above-mentioned probabilities behave as the underlying parameters and / or offspring distributions are allowed to vary. We include in this paper conjectures pertaining to the behaviour of the probability of draw in our game (including a phase transition phenomenon, in which the probability of draw goes from being $0$ to being strictly positive) as the parameter-pair $(p_{0},p_{1})$ is varied suitably while keeping the underlying offspring distribution of $T$ fixed.

math.PR

Towards Size-Independent Generalization Bounds for Deep Operator Nets

In recent times machine learning methods have made significant advances in becoming a useful tool for analyzing physical systems. A particularly active area in this theme has been "physics-informed machine learning" which focuses on using neural nets for numerically solving differential equations. In this work, we aim to advance the theory of measuring out-of-sample error while training DeepONets - which is among the most versatile ways to solve P.D.E systems in one-shot. Firstly, for a class of DeepONets, we prove a bound on their Rademacher complexity which does not explicitly scale with the width of the nets involved. Secondly, we use this to show how the Huber loss can be chosen so that for these DeepONet classes generalization error bounds can be obtained that have no explicit dependence on the size of the nets. The effective capacity measure for DeepONets that we thus derive is also shown to correlate with the behavior of generalization error in experiments.

cs.LG

GARCHX-NoVaS: A Model-free Approach to Incorporate Exogenous Variables

In this work, we explore the forecasting ability of a recently proposed normalizing and variance-stabilizing (NoVaS) transformation with the possible inclusion of exogenous variables. From an applied point-of-view, extra knowledge such as fundamentals- and sentiments-based information could be beneficial to improve the prediction accuracy of market volatility if they are incorporated into the forecasting process. In the classical approach, these models including exogenous variables are typically termed GARCHX-type models. Being a Model-free prediction method, NoVaS has generally shown more accurate, stable and robust (to misspecifications) performance than that compared to classical GARCH-type methods. This motivates us to extend this framework to the GARCHX forecasting as well. We derive the NoVaS transformation needed to include exogenous covariates and then construct the corresponding prediction procedure. We show through extensive simulation studies that bolster our claim that the NoVaS method outperforms traditional ones, especially for long-term time aggregated predictions. We also provide an interesting data analysis to exhibit how our method could possibly shed light on the role of geopolitical risks in forecasting volatility in national stock market indices for three different countries in Europe.

econ.EM

Gaussian Approximation For Non-stationary Time Series with Optimal Rate and Explicit Construction

Statistical inference for time series such as curve estimation for time-varying models or testing for existence of change-point have garnered significant attention. However, these works are generally restricted to the assumption of independence and/or stationarity at its best. The main obstacle is that the existing Gaussian approximation results for non-stationary processes only provide an existential proof and thus they are difficult to apply. In this paper, we provide two clear paths to construct such a Gaussian approximation for non-stationary series. While the first one is theoretically more natural, the second one is practically implementable. Our Gaussian approximation results are applicable for a very large class of non-stationary time series, obtain optimal rates and yet have good applicability. Building on such approximations, we also show theoretical results for change-point detection and simultaneous inference in presence of non-stationary errors. Finally we substantiate our theoretical results with simulation studies and real data analysis.

math.ST

Ergodicity of a generalized probabilistic cellular automaton with parity-based neighbourhoods

We study a one-dimensional generalized probabilistic cellular automaton $E_{p, q}$ with universe $\mathbb Z$, alphabet $\mathcal A = \{0, 1\}$, parameters $p$ and $q$ such that $0 < p+q \leq 1$ and two neighbourhoods $\mathcal N_0 = \{0, 1\}$ and $\mathcal N = \{1, 2\}$. The state $E_{p, q} η(x)$ of any $x \in \mathbb Z$ under the application of $E_{p, q}$ is a random variable whose probability distribution depends on the states $η(x + y)$ for $y \in \mathcal N_i$ where $i$ has the same parity as $x$. We establish ergodicity of this GPCA for various ranges of values of $p$ and $q$ via its connection with a suitable percolation game on a two-dimensional lattice. For these same ranges of values of $p$ and $q$, we show that the above-mentioned game has probability $0$ of resulting in a draw.

math.PR

Phase transition in percolation games on rooted Galton-Watson trees

We study the bond percolation game and the site percolation game on the rooted Galton-Watson tree $T_χ$ with offspring distribution $χ$. We obtain the probabilities of win, loss and draw for each player in terms of the fixed points of functions that involve the probability generating function $G$ of $χ$, and the parameters $p$ and $q$. Here, $p$ is the probability with which each edge (respectively vertex) of $T_χ$ is labeled a trap in the bond (respectively site) percolation game, and $q$ is the probability with which each edge (respectively vertex) of $T_χ$ is labeled a target in the bond (respectively site) percolation game. We obtain a necessary and sufficient condition for the probability of draw to be $0$ in each game, and we examine how this condition simplifies to yield very precise phase transition results when $χ$ is Binomial$(d,π)$, Poisson$(λ)$, or Negative Binomial$(r,π)$, or when $χ$ is supported on $\{0,d\}$ for some $d \in \mathbb{N}$, $d \geqslant 2$. It is fascinating to note that, while all other specific classes of offspring distributions we consider in this paper exhibit phase transition phenomena as the parameter-pair $(p,q)$ varies, the probability that the bond percolation game results in a draw remains $0$ for all values of $(p,q)$ when $χ$ is Geometric$(π)$, for all $0 < π\leqslant 1$. By establishing a connection between these games and certain finite state probabilistic tree automata on rooted $d$-regular trees, we obtain a precise description of the regime (in terms of $p$, $q$ and $d$) in which these automata exhibit ergodicity or weak spatial mixing.

math.PR

The spread of an epidemic: a game-theoretic approach

We introduce and study a model stemming from game theory for the spread of an epidemic throughout a given population. Each agent is allowed to choose an action whose value dictates to what extent they limit their social interactions, if at all. Each of them is endowed with a certain amount of immunity such that if the viral risk/exposure is more than that they get infected. We consider a discrete-time stochastic process where, at the beginning of each epoch, a randomly chosen agent is allowed to update their action, which they do with the aim of maximizing a utility function that is a function of the state in which the process is currently in. The state itself is determined by the subset of infected agents at the beginning of that epoch, and the most recent action profile of all the agents. Our main results are concerned with the limiting distributions of both the cardinality of the subset of infected agents and the action profile as time approaches infinity, considered under various settings (such as the initial action profile we begin with, the value of each agent's immunity etc.). We also provide some simulations to show that the final asymptotic distributions for the cardinality of infected set are almost always achieved within the first few epochs.

math.PR

An NLP-Assisted Bayesian Time Series Analysis for Prevalence of Twitter Cyberbullying During the COVID-19 Pandemic

COVID-19 has brought about many changes in social dynamics. Stay-at-home orders and disruptions in school teaching can influence bullying behavior in-person and online, both of which leading to negative outcomes in victims. To study cyberbullying specifically, 1 million tweets containing keywords associated with abuse were collected from the beginning of 2019 to the end of 2021 with the Twitter API search endpoint. A natural language processing model pre-trained on a Twitter corpus generated probabilities for the tweets being offensive and hateful. To overcome limitations of sampling, data was also collected using the count endpoint. The fraction of tweets from a given daily sample marked as abusive is multiplied to the number reported by the count endpoint. Once these adjusted counts are assembled, a Bayesian autoregressive Poisson model allows one to study the mean trend and lag functions of the data and how they vary over time. The results reveal strong weekly and yearly seasonality in hateful speech but with slight differences across years that may be attributed to COVID-19.

cs.SI

On a class of probabilistic cellular automata with size-$3$ neighbourhood and their applications in percolation games

Different versions of percolation games on $\mathbb{Z}^{2}$, with parameters $p$ and $q$ that indicate, respectively, the probability with which a site in $\mathbb{Z}^{2}$ is labeled a trap and the probability with which it is labeled a target, are shown to have probability $0$ of culminating in draws when $p+q > 0$. We show that, for fixed $p$ and $q$, the probability of draw in each of these games is $0$ if and only if a certain $1$-dimensional probabilistic cellular automaton (PCA) $F_{p,q}$ with a size-$3$ neighbourhood is ergodic. This allows us to conclude that $F_{p,q}$ is ergodic whenever $p+q > 0$, thereby rigorously establishing ergodicity for a considerable class of PCAs.

math.PR

Depth-2 Neural Networks Under a Data-Poisoning Attack

In this work, we study the possibility of defending against data-poisoning attacks while training a shallow neural network in a regression setup. We focus on doing supervised learning for a class of depth-2 finite-width neural networks, which includes single-filter convolutional networks. In this class of networks, we attempt to learn the network weights in the presence of a malicious oracle doing stochastic, bounded and additive adversarial distortions on the true output during training. For the non-gradient stochastic algorithm that we construct, we prove worst-case near-optimal trade-offs among the magnitude of the adversarial attack, the weight approximation accuracy, and the confidence achieved by the proposed algorithm. As our algorithm uses mini-batching, we analyze how the mini-batch size affects convergence. We also show how to utilize the scaling of the outer layer weights to counter output-poisoning attacks depending on the probability of attack. Lastly, we give experimental evidence demonstrating how our algorithm outperforms stochastic gradient descent under different input data distributions, including instances of heavy-tailed distributions.

cs.LG

An Empirical Study of the Occurrence of Heavy-Tails in Training a ReLU Gate

A particular direction of recent advance about stochastic deep-learning algorithms has been about uncovering a rather mysterious heavy-tailed nature of the stationary distribution of these algorithms, even when the data distribution is not so. Moreover, the heavy-tail index is known to show interesting dependence on the input dimension of the net, the mini-batch size and the step size of the algorithm. In this short note, we undertake an experimental study of this index for S.G.D. while training a $\relu$ gate (in the realizable and in the binary classification setup) and for a variant of S.G.D. that was proven in Karmakar and Mukherjee (2022) for ReLU realizable data. From our experiments we conjecture that these two algorithms have similar heavy-tail behaviour on any data where the latter can be proven to converge. Secondly, we demonstrate that the heavy-tail index of the late time iterates in this model scenario has strikingly different properties than either what has been proven for linear hypothesis classes or what has been previously demonstrated for large nets.

cs.LG