Searcharxiv⌕ Search

arXiv subjects

Soumendu Sundar Mukherjee

Publications and source records attributed to Soumendu Sundar Mukherjee.

At least 19 recordsLinked to original sources

A statistical approach to bias in zero-shot learning: the lens of handwriting recognition

Generalized zero-shot learning (GZSL) has emerged as an important paradigm for visual recognition systems that must generalize to classes that were not observed during training. Traditional GZSL techniques are limited by their applicability to a relatively small number of such unseen classes, scalability beyond which is challenging due to its well-known misclassification bias towards classes observed during training. In this work, we investigate the GZSL paradigm through the lens of zero-shot handwritten word recognition over extremely large vocabularies. We propose a statistical approach to rectifying this bias, which views any classical GZSL feature learner as a black box mechanism whose intrinsic bias in identifying the training status (seen vs. unseen) of a typical data point we aim to correct, similar to an out of distribution inferential problem. Our method leverages a simple two-stage hierarchical architecture, combining a classical GZSL blackbox in the first stage and an ensemble of lightweight Monte Carlo bias-correctors in the second. Once debiased, the classification of test data is undertaken only restricted to its predicted training status via well-founded statistical methods (eg nearest neighbour, logistic regression and random forests). We achieve relative accuracy improvements of over 20% in the classification of unseen words compared to established techniques. A key outcome is that word recognition over large scale vocabularies is amenable to a much lower dimensional representation (~15 dimensions). Our approach is underpinned by mathematical analysis that captures the essence of the statistical approach to bias correction. Our approach to bias rectification can be combined in a turn-key fashion with any classical GZSL learner as a blackbox, thereby suggesting a wide scope of applicability of this method for a wide variety of GZSL implementations in different domains.

stat.ML↗

Approximation Theory for Neural Networks: Old and New

Universal approximation theorems provide a mathematical explanation for the expressive power of neural networks. They assert that, under mild conditions on the activation function, feedforward neural networks are dense in broad function classes, such as continuous functions on compact subsets of $\mathbb{R}^d$, $L^p$ spaces, or Sobolev spaces. Over the past four decades, these qualitative universality results have evolved into a rich quantitative theory addressing approximation rates, parameter efficiency, and the role of architectural features such as depth and width. This survey presents several glimpses into this theory. We review classical density results for single-hidden-layer networks, as well as quantitative bounds that relate approximation error to network size and smoothness assumptions on target functions. Particular emphasis is placed on depth--width trade-offs and on results demonstrating that deeper architectures can achieve superior parameter efficiency for structured function classes. In addition to standard feedforward neural networks, we also review recent developments on Kolmogorov--Arnold Networks (KANs), which offer an alternative architectural paradigm and whose approximation-theoretic properties have begun to attract significant theoretical attention.

cs.LG↗

Elephant random walks on infinite Cayley trees

We introduce a generalisation of Schütz and Trimper's elephant random walk to finitely generated groups. We focus on the simplest non-abelian setting, i.e. groups whose Cayley graphs are homogeneous trees of degree $d \ge 3$. We show that the asymptotic speed of the walk does not depend on the memory parameter $p \in [0, 1)$ and equals $\frac{d - 2}{d}$, the asymptotic speed of simple random walk on these graphs. We also establish upper bounds on the rate of convergence to the limiting speed. These upper bounds depend on $p$ and exhibit a phase transition at the critical value $p_d = \frac{d + 1}{2d}$. Numerical experiments suggest that these upper bounds are tight. Along the way, we also obtain estimates on the return probability.

math.PR↗

Elephant random walk on the infinite dihedral group $\mathbb{Z}_2 * \mathbb{Z}_2$

Elephant random walks were studied recently in \cite{mukherjee2025elephant} on the groups $\mathbb{Z}^{*d_1} * \mathbb{Z}_2^{*d_2}$ whose Cayley graphs are infinite $d$-regular trees with $d = 2d_1 + d_2$. It was found that for $d \ge 3$, the elephant walk is ballistic with the same asymptotic speed $\frac{d - 2}{d}$ as the simple random walk and the memory parameter appears only in the rate of convergence to the limiting speed. In the $d = 2$ case, there are two such groups, both having the bi-infinite path as their Cayley graph. For $(d_1, d_2) = (1, 0)$, the walk is the usual elephant random walk on $\mathbb{Z}$, which exhibits anomalous diffusion. In this article, we study the other case, namely $(d_1, d_2) = (0, 2)$, which corresponds to the infinite dihedral group $D_\infty \cong \mathbb{Z}_2 * \mathbb{Z}_2$. Unlike the classical ERW on $\mathbb{Z}$, which is a time-inhomogeneous Markov chain, the ERW on $D_{\infty}$ is non-Markovian. We show that the first and second order behaviours of the \emph{signed location} of the walker agree with those of the simple symmetric random walk on $\mathbb{Z}$, with the memory parameter essentially manifesting itself via a lower order correction term that can be written as an explicit functional of the elephant walk on $\mathbb{Z}$. Our result demonstrates that unlike the simple random walk, the elephant walk is sensitive to local algebraic relations. Indeed, although $D_{\infty}$ is virtually abelian, containing $\mathbb{Z}$ as a finite-index subgroup, the involutive nature of its generators effectively neutralises memory, thereby ruling out any potential superdiffusive behaviour, in contrast to the superdiffusion observed on its abelian cousin $\mathbb{Z}$.

math.PR↗

Filtering through a topological lens: homology for point processes on the time-frequency plane

We introduce a very general approach to the analysis of signals from their noisy measurements from the perspective of Topological Data Analysis (TDA). While TDA has emerged as a powerful analytical tool for data with pronounced topological structures, here we demonstrate its applicability for general problems of signal processing, without any a-priori geometric feature. Our methods are well-suited to a wide array of time-dependent signals in different scientific domains, with acoustic signals being a particularly important application. We invoke time-frequency representations of such signals, focusing on their zeros which are gaining salience as a signal processing tool in view of their stability properties. Leveraging state-of-the-art topological concepts, such as stable and minimal volumes, we develop a complete suite of TDA-based methods to explore the delicate stochastic geometry of these zeros, capturing signals based on the disruption they cause to this rigid, hyperuniform spatial structure. Unlike classical spatial data tools, TDA is able to capture the full spectrum of the stochastic geometry of the zeros, thereby leading to powerful inferential outcomes that are underpinned by a principled statistical foundation. This is reflected in the power and versatility of our applications, which include competitive performance in processing. a wide variety of audio signals (esp. in low SNR regimes), effective detection and reconstruction of gravitational wave signals (a reputed signal processing challenge with non-Gaussian noise), and medical time series data from EEGs, indicating a wide horizon for the approach and methods introduced in this paper.

eess.SP↗

Learning under Latent Group Sparsity via Diffusion on Networks

Group or cluster structure on explanatory variables in machine learning problems is a very general phenomenon, which has attracted broad interest from practitioners and theoreticians alike. In this work we contribute an approach to sparse learning under such group structure, that does not require prior information on the group identities. Our paradigm is motivated by the Laplacian geometry of an underlying network with a related community structure, and proceeds by directly incorporating this into a penalty that is effectively computed via a heat-flow-based local network dynamics. The proposed penalty interpolates between the lasso and the group lasso penalties, the runtime of the heat-flow dynamics being the interpolating parameter. As such it can automatically default to lasso when the group structure reflected in the Laplacian is weak. In fact, we demonstrate a data-driven procedure to construct such a network based on the available data. Notably, we dispense with computationally intensive pre-processing involving clustering of variables, spectral or otherwise. Our technique is underpinned by rigorous theorems that guarantee its effective performance and provide bounds on its sample complexity. In particular, in a wide range of settings, it provably suffices to run the diffusion for time that is only logarithmic in the problem dimensions. We explore in detail the interfaces of our approach with key statistical physics models in network science, such as the Gaussian Free Field and the Stochastic Block Model. Our work raises the possibility of applying similar diffusion-based techniques to classical learning tasks, exploiting the interplay between geometric, dynamical and stochastic structures underlying the data.

stat.ML↗

A new approach to locally adaptive polynomial regression

Adaptive bandwidth selection is a fundamental challenge in nonparametric regression. This paper introduces a new bandwidth selection procedure inspired by the optimality criteria for $\ell_0$-penalized regression. Although similar in spirit to Lepski's method and its variants in selecting the largest interval satisfying an admissibility criterion, our approach stems from a distinct philosophy, utilizing criteria based on $\ell_2$-norms of interval projections rather than explicit point and variance estimates. We obtain non-asymptotic risk bounds for the local polynomial regression methods based on our bandwidth selection procedure which adapt (near-)optimally to the local Hölder exponent of the underlying regression function simultaneously at all points in its domain. Furthermore, we show that there is a single ideal choice of a global tuning parameter in each case under which the above-mentioned local adaptivity holds. The optimal risks of our methods derive from the properties of solutions to a new ``bandwidth selection equation'' which is of independent interest. We believe that the principles underlying our approach provide a new perspective to the classical yet ever relevant problem of locally adaptive nonparametric regression.

stat.ML↗

Spectra of contractions of the Gaussian Orthogonal Tensor Ensemble

In this article, we study the spectra of matrix-valued contractions of the Gaussian Orthogonal Tensor Ensemble (GOTE). Let $\mathcal{G}$ denote a random tensor of order $r$ and dimension $n$ drawn from the density \[ f(\mathcal{G}) \propto \exp\bigg(-\frac{1}{2r}\|\mathcal{G}\|^2_{\mathrm{F}}\bigg). \] For $\mathbf{w} \in \mathbb{S}^{n - 1}$, the unit-sphere in $\mathbb{R}^n$, we consider the matrix-valued contraction $\mathcal{G} \cdot \mathbf{w}^{\otimes (r - 2)}$ when both $r$ and $n$ go to infinity such that $r / n \to c \in [0, \infty]$. We obtain semi-circle bulk-limits in all regimes, generalising the works of Goulart et al. (2022); Au and Garza-Vargas (2023); Bonnin (2024) in the fixed-$r$ setting. We also study the edge-spectrum. We obtain a Baik-Ben Arous-Péché phase-transition for the largest and the smallest eigenvalues at $r = 4$, generalising a result of Mukherjee et al. (2024) in the context of adjacency matrices of random hypergraphs. We also show that the extreme eigenvectors of $\mathcal{G} \cdot \mathbf{w}^{\otimes (r - 2)}$ contain non-trivial information about the contraction direction $\mathbf{w}$. Finally, we report some results, in the case $r = 4$, on mixed contractions $\mathcal{G} \cdot \mathbf{u} \otimes \mathbf{v}$, $\mathbf{u}, \mathbf{v} \in \mathbb{S}^{n - 1}$. While the total variation distance between the joint distribution of the entries of $\mathcal{G} \cdot \mathbf{u} \otimes \mathbf{v}$ and that of $\mathcal{G} \cdot \mathbf{u} \otimes \mathbf{u}$ goes to $0$ when $\|\mathbf{u} - \mathbf{v}\| = o(n^{-1})$, the bulk and the largest eigenvalues of these two matrices have the same limit profile as long as $\|\mathbf{u} - \mathbf{v}\| = o(1)$. Furthermore, it turns out that there are no outlier eigenvalues in the spectrum of $\mathcal{G} \cdot \mathbf{u} \otimes \mathbf{v}$ when $\langle \mathbf{u}, \mathbf{v} \rangle = o(1)$.

math.PR↗

Optimal Transfer Learning for Missing Not-at-Random Matrix Completion

We study transfer learning for matrix completion in a Missing Not-at-Random (MNAR) setting that is motivated by biological problems. The target matrix $Q$ has entire rows and columns missing, making estimation impossible without side information. To address this, we use a noisy and incomplete source matrix $P$, which relates to $Q$ via a feature shift in latent space. We consider both the active and passive sampling of rows and columns. We establish minimax lower bounds for entrywise estimation error in each setting. Our computationally efficient estimation framework achieves this lower bound for the active setting, which leverages the source data to query the most informative rows and columns of $Q$. This avoids the need for incoherence assumptions required for rate optimality in the passive sampling setting. We demonstrate the effectiveness of our approach through comparisons with existing algorithms on real-world biological datasets.

cs.LG↗

Construction of product $*$-probability spaces via free cumulants

It is well known that free independence is equivalent to the vanishing of mixed free cumulants. The purpose of this short note is to build free products of $*$-probability spaces using this as the definition of freeness and relying on free cumulants instead of moments.

math.OA↗

Consistent model selection in the spiked Wigner model via AIC-type criteria

Consider the spiked Wigner model \[ X = \sum_{i = 1}^k λ_i u_i u_i^\top + σG, \] where $G$ is an $N \times N$ GOE random matrix, and the eigenvalues $λ_i$ are all spiked, i.e. above the Baik-Ben Arous-Péché (BBP) threshold $σ$. We consider AIC-type model selection criteria of the form \[ -2 \, (\text{maximised log-likelihood}) + γ\, (\text{number of parameters}) \] for estimating the number $k$ of spikes. For $γ> 2$, the above criterion is strongly consistent provided $λ_k > λ_γ$, where $λ_γ$ is a threshold strictly above the BBP threshold, whereas for $γ< 2$, it almost surely overestimates $k$. Although AIC (which corresponds to $γ= 2$) is not strongly consistent, we show that taking $γ= 2 + δ_N$, where $δ_N \to 0$ and $δ_N \gg N^{-2/3}$, results in a weakly consistent estimator of $k$. We further show that a soft minimiser of AIC, where one chooses the least complex model whose AIC score is close to the minimum AIC score, is strongly consistent. Based on a spiked (generalised) Wigner representation, we also develop similar model selection criteria for consistently estimating the number of communities in a balanced stochastic block model under some sparsity restrictions.

math.ST↗

Edge spectra of Gaussian random symmetric matrices with correlated entries

We study the largest eigenvalue of a Gaussian random symmetric matrix $X_n$, with zero-mean, unit variance entries satisfying the condition $\sup_{(i, j) \ne (i', j')}|\mathbb{E}[X_{ij} X_{i'j'}]| = O(n^{-(1 + \varepsilon)})$, where $\varepsilon > 0$. It follows from Catalano et al. (2024) that the empirical spectral distribution of $n^{-1/2} X_n$ converges weakly almost surely to the standard semi-circle law. Using a Füredi-Komlós-type high moment analysis, we show that the largest eigenvalue $λ_1(n^{-1/2} X_n)$ of $n^{-1/2} X_n$ converges almost surely to $2$. This result is essentially optimal in the sense that one cannot take $\varepsilon = 0$ and still obtain an almost sure limit of $2$. We also derive Gaussian fluctuation results for the largest eigenvalue in the case where the entries have a common non-zero mean. Let $Y_n = X_n + \fracλ{\sqrt{n}}\mathbf{1} \mathbf{1}^\top$. When $\varepsilon \ge 1$ and $λ\gg n^{1/4}$, we show that \[ n^{1/2}\bigg(λ_1(n^{-1/2} Y_n) - λ- \frac{1}λ\bigg) \xrightarrow{d} \sqrt{2} Z, \] where $Z$ is a standard Gaussian. On the other hand, when $0 < \varepsilon < 1$, we have $\mathrm{Var}(\frac{1}{n}\sum_{i, j}X_{ij}) = O(n^{1 - \varepsilon})$. Assuming that $\mathrm{Var}(\frac{1}{n}\sum_{i, j} X_{ij}) = σ^2 n^{1 - \varepsilon} (1 + o(1))$, if $λ\gg n^{\varepsilon/4}$, then we have \[ n^{\varepsilon/2}\bigg(λ_1(n^{-1/2} Y_n) - λ- \frac{1}λ\bigg) \xrightarrow{d} σZ. \] While the ranges of $λ$ in these fluctuation results are certainly not optimal, a striking aspect is that different scalings are required in the two regimes $0 < \varepsilon < 1$ and $\varepsilon \ge 1$.

math.PR↗

Spectra of adjacency and Laplacian matrices of Erdős-Rényi hypergraphs

We study adjacency and Laplacian matrices of Erdős-Rényi $r$-uniform hypergraphs on $n$ vertices with hyperedge inclusion probability $p$, in the setting where $r$ can vary with $n$ such that $r / n \to c \in [0, 1)$. Adjacency matrices of hypergraphs are contractions of adjacency tensors and their entries exhibit long range correlations. We show that under the Erdős-Rényi model, the expected empirical spectral distribution of an appropriately normalised hypergraph adjacency matrix converges weakly to the semi-circle law with variance $(1 - c)^2$ as long as $\frac{d_{\avg}}{r^7} \to \infty$, where $d_{\avg} = \binom{n-1}{r-1} p$. In contrast with the Erdős-Rényi random graph ($r = 2$), two eigenvalues stick out of the bulk of the spectrum. When $r$ is fixed and $d_{\avg} \gg n^{r - 2} \log^4 n$, we uncover an interesting Baik-Ben Arous-Péché (BBP) phase transition at the value $r = 3$. For $r \in \{2, 3\}$, an appropriately scaled largest (resp. smallest) eigenvalue converges in probability to $2$ (resp. $-2$), the right (resp. left) end point of the support of the standard semi-circle law, and when $r \ge 4$, it converges to $\sqrt{r - 2} + \frac{1}{\sqrt{r - 2}}$ (resp. $-\sqrt{r - 2} - \frac{1}{\sqrt{r - 2}}$). Further, in a Gaussian version of the model we show that an appropriately scaled largest (resp. smallest) eigenvalue converges in distribution to $\frac{c}{2} ζ+ \big[\frac{c^2}{4}ζ^2 + c(1 - c)\big]^{1/2}$ (resp. $\frac{c}{2} ζ- \big[\frac{c^2}{4}ζ^2 + c(1 - c)\big]^{1/2}$), where $ζ$ is a standard Gaussian. We also establish analogous results for the bulk and edge eigenvalues of the associated Laplacian matrices.

math.PR↗

Bulk Spectra of Truncated Sample Covariance Matrices

Determinantal Point Processes (DPPs), which originate from quantum and statistical physics, are known for modelling diversity. Recent research [Ghosh and Rigollet (2020)] has demonstrated that certain matrix-valued $U$-statistics (that are truncated versions of the usual sample covariance matrix) can effectively estimate parameters in the context of Gaussian DPPs and enhance dimension reduction techniques, outperforming standard methods like PCA in clustering applications. This paper explores the spectral properties of these matrix-valued $U$-statistics in the null setting of an isotropic design. These matrices may be represented as $X L X^\top$, where $X$ is a data matrix and $L$ is the Laplacian matrix of a random geometric graph associated to $X$. The main mathematically interesting twist here is that the matrix $L$ is dependent on $X$. We give complete descriptions of the bulk spectra of these matrix-valued $U$-statistics in terms of the Stieltjes transforms of their empirical spectral measures. The results and the techniques are in fact able to address a broader class of kernelised random matrices, connecting their limiting spectra to generalised Marčenko-Pastur laws and free probability.

math.ST↗

Transfer Learning for Latent Variable Network Models

We study transfer learning for estimation in latent variable network models. In our setting, the conditional edge probability matrices given the latent variables are represented by $P$ for the source and $Q$ for the target. We wish to estimate $Q$ given two kinds of data: (1) edge data from a subgraph induced by an $o(1)$ fraction of the nodes of $Q$, and (2) edge data from all of $P$. If the source $P$ has no relation to the target $Q$, the estimation error must be $Ω(1)$. However, we show that if the latent variables are shared, then vanishing error is possible. We give an efficient algorithm that utilizes the ordering of a suitably defined graph distance. Our algorithm achieves $o(1)$ error and does not assume a parametric form on the source or target networks. Next, for the specific case of Stochastic Block Models we prove a minimax lower bound and show that a simple algorithm achieves this rate. Finally, we empirically demonstrate our algorithm's use on real-world and simulated graph transfer problems.

cs.LG↗

Implicit Regularization via Spectral Neural Networks and Non-linear Matrix Sensing

The phenomenon of implicit regularization has attracted interest in recent years as a fundamental aspect of the remarkable generalizing ability of neural networks. In a nutshell, it entails that gradient descent dynamics in many neural nets, even without any explicit regularizer in the loss function, converges to the solution of a regularized learning problem. However, known results attempting to theoretically explain this phenomenon focus overwhelmingly on the setting of linear neural nets, and the simplicity of the linear structure is particularly crucial to existing arguments. In this paper, we explore this problem in the context of more realistic neural networks with a general class of non-linear activation functions, and rigorously demonstrate the implicit regularization phenomenon for such networks in the setting of matrix sensing problems, together with rigorous rate guarantees that ensure exponentially fast convergence of gradient descent.In this vein, we contribute a network architecture called Spectral Neural Networks (abbrv. SNN) that is particularly suitable for matrix learning problems. Conceptually, this entails coordinatizing the space of matrices by their singular values and singular vectors, as opposed to by their entries, a potentially fruitful perspective for matrix learning. We demonstrate that the SNN architecture is inherently much more amenable to theoretical analysis than vanilla neural nets and confirm its effectiveness in the context of matrix sensing, via both mathematical guarantees and empirical investigations. We believe that the SNN architecture has the potential to be of wide applicability in a broad class of matrix learning scenarios.

cs.LG↗

The "visible" Wigner matrix

We consider the ``visible'' Wigner matrix, a Wigner matrix whose $(i, j)$-th entry is coerced to zero if $i, j$ are co-prime. Using a recent result from elementary number theory on co-primality patterns in integers, we show that the limiting spectral distribution of this matrix exists, and give explicit descriptions of its moments in terms of infinite products over primes $p$ of certain polynomials evaluated at $1/p$. We also consider the complementary ``invisible'' Wigner matrix.

math.PR↗

Minimax-optimal estimation for sparse multi-reference alignment with collision-free signals

The Multi-Reference Alignment (MRA) problem aims at the recovery of an unknown signal from repeated observations under the latent action of a group of cyclic isometries, in the presence of additive noise of high intensity $σ$. It is a more tractable version of the celebrated cryo EM model. In the crucial high noise regime, it is known that its sample complexity scales as $σ^6$. Recent investigations have shown that for the practically significant setting of sparse signals, the sample complexity of the maximum likelihood estimator asymptotically scales with the noise level as $σ^4$. In this work, we investigate minimax optimality for signal estimation under the MRA model for so-called collision-free signals. In particular, this signal class covers the setting of generic signals of dilute sparsity (wherein the support size $s=O(L^{1/3})$, where $L$ is the ambient dimension. We demonstrate that the minimax optimal rate of estimation in for the sparse MRA problem in this setting is $σ^2/\sqrt{n}$, where $n$ is the sample size. In particular, this widely generalizes the sample complexity asymptotics for the restricted MLE in this setting, establishing it as the statistically optimal estimator. Finally, we demonstrate a concentration inequality for the restricted MLE on its deviations from the ground truth.

math.ST↗