SearcharxivSearch

arXiv subjects

Subhro Ghosh

Publications and source records attributed to Subhro Ghosh.

At least 19 recordsLinked to original sources

On the Statistical Optimality of Optimal Decision Trees

While globally optimal empirical risk minimization (ERM) decision trees have become computationally feasible and empirically successful, rigorous theoretical guarantees for their statistical performance remain limited. In this work, we develop a comprehensive statistical theory for ERM trees under random design in both high-dimensional regression and classification. We first establish sharp oracle inequalities that bound the excess risk of the ERM estimator relative to the best possible approximation achievable by any tree with at most $L$ leaves, thereby characterizing the interpretability-accuracy trade-off. We derive these results using a novel uniform concentration framework based on empirically localized Rademacher complexity. Furthermore, we derive minimax optimal rates over a novel function class: the piecewise sparse heterogeneous anisotropic Besov (PSHAB) space. This space explicitly captures three key structural features encountered in practice: sparsity, anisotropic smoothness, and spatial heterogeneity. While our main results are established under sub-Gaussianity, we also provide robust guarantees that hold under heavy-tailed noise settings. Together, these findings provide a principled foundation for the optimality of ERM trees and introduce empirical process tools broadly applicable to other highly adaptive, data-driven procedures.

stat.ML

Maximum entropy based testing in network models: ERGMs and constrained optimization

Stochastic network models play a central role across a wide range of scientific disciplines, and questions of statistical inference arise naturally in this context. In this paper we investigate goodness-of-fit and two-sample testing procedures for statistical networks based on the principle of maximum entropy (MaxEnt). Our approach formulates a constrained entropy-maximization problem on the space of networks, subject to prescribed structural constraints. The resulting test statistics are defined through the Lagrange multipliers associated with the constrained optimization problem, which, to our knowledge, is novel in the statistical networks literature. We establish consistency in the classical regime where the number of vertices is fixed. We then consider asymptotic regimes in which the graph size grows with the sample size, developing tests for both dense and sparse settings. In the dense case, we analyze exponential random graph models (ERGM) (including the Erd\"os-R\`enyi models), while in the sparse regime our theory applies to Erd{\"o}s-R{\`e}nyi graphs. Our analysis leverages recent advances in nonlinear large deviation theory for random graphs. We further show that the proposed Lagrange-multiplier framework connects naturally to classical score tests for constrained maximum likelihood estimation. The results provide a unified entropy-based framework for network model assessment across diverse growth regimes.

math.ST

Implicit Regularization via Spectral Neural Networks and Non-linear Matrix Sensing

The phenomenon of implicit regularization has attracted interest in recent years as a fundamental aspect of the remarkable generalizing ability of neural networks. In a nutshell, it entails that gradient descent dynamics in many neural nets, even without any explicit regularizer in the loss function, converges to the solution of a regularized learning problem. However, known results attempting to theoretically explain this phenomenon focus overwhelmingly on the setting of linear neural nets, and the simplicity of the linear structure is particularly crucial to existing arguments. In this paper, we explore this problem in the context of more realistic neural networks with a general class of non-linear activation functions, and rigorously demonstrate the implicit regularization phenomenon for such networks in the setting of matrix sensing problems, together with rigorous rate guarantees that ensure exponentially fast convergence of gradient descent.In this vein, we contribute a network architecture called Spectral Neural Networks (abbrv. SNN) that is particularly suitable for matrix learning problems. Conceptually, this entails coordinatizing the space of matrices by their singular values and singular vectors, as opposed to by their entries, a potentially fruitful perspective for matrix learning. We demonstrate that the SNN architecture is inherently much more amenable to theoretical analysis than vanilla neural nets and confirm its effectiveness in the context of matrix sensing, via both mathematical guarantees and empirical investigations. We believe that the SNN architecture has the potential to be of wide applicability in a broad class of matrix learning scenarios.

cs.LG

Minimax-optimal estimation for sparse multi-reference alignment with collision-free signals

The Multi-Reference Alignment (MRA) problem aims at the recovery of an unknown signal from repeated observations under the latent action of a group of cyclic isometries, in the presence of additive noise of high intensity $\sigma$. It is a more tractable version of the celebrated cryo EM model. In the crucial high noise regime, it is known that its sample complexity scales as $\sigma^6$. Recent investigations have shown that for the practically significant setting of sparse signals, the sample complexity of the maximum likelihood estimator asymptotically scales with the noise level as $\sigma^4$. In this work, we investigate minimax optimality for signal estimation under the MRA model for so-called collision-free signals. In particular, this signal class covers the setting of generic signals of dilute sparsity (wherein the support size $s=O(L^{1/3})$, where $L$ is the ambient dimension. We demonstrate that the minimax optimal rate of estimation in for the sparse MRA problem in this setting is $\sigma^2/\sqrt{n}$, where $n$ is the sample size. In particular, this widely generalizes the sample complexity asymptotics for the restricted MLE in this setting, establishing it as the statistically optimal estimator. Finally, we demonstrate a concentration inequality for the restricted MLE on its deviations from the ground truth.

math.ST

Learning Networks from Gaussian Graphical Models and Gaussian Free Fields

We investigate the problem of estimating the structure of a weighted network from repeated measurements of a Gaussian Graphical Model (GGM) on the network. In this vein, we consider GGMs whose covariance structures align with the geometry of the weighted network on which they are based. Such GGMs have been of longstanding interest in statistical physics, and are referred to as the Gaussian Free Field (GFF). In recent years, they have attracted considerable interest in the machine learning and theoretical computer science. In this work, we propose a novel estimator for the weighted network (equivalently, its Laplacian) from repeated measurements of a GFF on the network, based on the Fourier analytic properties of the Gaussian distribution. In this pursuit, our approach exploits complex-valued statistics constructed from observed data, that are of interest on their own right. We demonstrate the effectiveness of our estimator with concrete recovery guarantees and bounds on the required sample complexity. In particular, we show that the proposed statistic achieves the parametric rate of estimation for fixed network size. In the setting of networks growing with sample size, our results show that for Erdos-Renyi random graphs $G(d,p)$ above the connectivity threshold, we demonstrate that network recovery takes place with high probability as soon as the sample size $n$ satisfies $n \gg d^4 \log d \cdot p^{-2}$.

math.ST

Approximate Gibbsian structure in strongly correlated point fields and generalized Gaussian zero ensembles

Gibbsian structure in random point fields has been a classical tool for studying their spatial properties. However, exact Gibbs property is available only in a relatively limited class of models, and it does not adequately address many random fields with a strongly dependent spatial structure. In this work, we provide a general framework for approximate Gibbsian structure for strongly correlated random point fields. These include processes that exhibit strong spatial rigidity, in particular, a certain one-parameter family of analytic Gaussian zero point fields, namely the $\alpha$-GAFs. Our framework entails conditions that may be verified via finite particle approximations to the process, a phenomenon that we call an approximate Gibbs property. We show that these enable one to compare the spatial conditional measures in the infinite volume limit with Gibbs-type densities supported on appropriate singular manifolds, a phenomenon we refer to as a generalized Gibbs property. We demonstrate the scope of our approach by showing that a generalized Gibbs property holds with a logarithmic pair potential for the $\alpha$-GAFs for any value of $\alpha$. This establishes the level of rigidity of the $\alpha$-GAF zero process to be exactly $\lfloor \frac{1}{\alpha} \rfloor$, settling in the affirmative an open question regarding the existence of point processes with any specified level of rigidity. For processes such as the zeros of $\alpha$-GAFs, which involve complex, many-body interactions, our results imply that the local behaviour of the random points still exhibits 2D Coulomb-type repulsion in the short range. Our techniques can be leveraged to estimate the relative energies of configurations under local perturbations, with possible implications for dynamics and stochastic geometry on strongly correlated random point fields.

math.PR

Determinantal point processes based on orthogonal polynomials for sampling minibatches in SGD

Stochastic gradient descent (SGD) is a cornerstone of machine learning. When the number N of data items is large, SGD relies on constructing an unbiased estimator of the gradient of the empirical risk using a small subset of the original dataset, called a minibatch. Default minibatch construction involves uniformly sampling a subset of the desired size, but alternatives have been explored for variance reduction. In particular, experimental evidence suggests drawing minibatches from determinantal point processes (DPPs), distributions over minibatches that favour diversity among selected items. However, like in recent work on DPPs for coresets, providing a systematic and principled understanding of how and why DPPs help has been difficult. In this work, we contribute an orthogonal polynomial-based DPP paradigm for minibatch sampling in SGD. Our approach leverages the specific data distribution at hand, which endows it with greater sensitivity and power over existing data-agnostic methods. We substantiate our method via a detailed theoretical analysis of its convergence properties, interweaving between the discrete data set and the underlying continuous domain. In particular, we show how specific DPPs and a string of controlled approximations can lead to gradient estimators with a variance that decays faster with the batchsize than under uniform sampling. Coupled with existing finite-time guarantees for SGD on convex objectives, this entails that, DPP minibatches lead to a smaller bound on the mean square approximation error than uniform minibatches. Moreover, our estimators are amenable to a recent algorithm that directly samples linear statistics of DPPs (i.e., the gradient estimator) without sampling the underlying DPP (i.e., the minibatch), thereby reducing computational overhead. We provide detailed synthetic as well as real data experiments to substantiate our theoretical claims.

stat.ML

Gaussian Determinantal Processes: a new model for directionality in data

Determinantal point processes (a.k.a. DPPs) have recently become popular tools for modeling the phenomenon of negative dependence, or repulsion, in data. However, our understanding of an analogue of a classical parametric statistical theory is rather limited for this class of models. In this work, we investigate a parametric family of Gaussian DPPs with a clearly interpretable effect of parametric modulation on the observed points. We show that parameter modulation impacts the observed points by introducing directionality in their repulsion structure, and the principal directions correspond to the directions of maximal (i.e. the most long ranged) dependency. This model readily yields a novel and viable alternative to Principal Component Analysis (PCA) as a dimension reduction tool that favors directions along which the data is most spread out. This methodological contribution is complemented by a statistical analysis of a spiked model similar to that employed for covariance matrices as a framework to study PCA. These theoretical investigations unveil intriguing questions for further examination in random matrix theory, stochastic geometry and related topics.

stat.ML

Sparse Multi-Reference Alignment : Phase Retrieval, Uniform Uncertainty Principles and the Beltway Problem

Motivated by cutting-edge applications like cryo-electron microscopy (cryo-EM), the Multi-Reference Alignment (MRA) model entails the learning of an unknown signal from repeated measurements of its images under the latent action of a group of isometries and additive noise of magnitude $\sigma$. Despite significant interest, a clear picture for understanding rates of estimation in this model has emerged only recently, particularly in the high-noise regime $\sigma \gg 1$ that is highly relevant in applications. Recent investigations have revealed a remarkable asymptotic sample complexity of order $\sigma^6$ for certain signals whose Fourier transforms have full support, in stark contrast to the traditional $\sigma^2$ that arise in regular models. Often prohibitively large in practice, these results have prompted the investigation of variations around the MRA model where better sample complexity may be achieved. In this paper, we show that sparse signals exhibit an intermediate $\sigma^4$ sample complexity even in the classical MRA model. Further, we characterise the dependence of the estimation rate on the support size $s$ as $O_p(1)$ and $O_p(s^{3.5})$ in the dilute and moderate regimes of sparsity respectively. Our techniques have implications for the problem of crystallographic phase retrieval, indicating a certain local uniqueness for the recovery of sparse signals from their power spectrum. Our results explore and exploit connections of the MRA estimation problem with two classical topics in applied mathematics: the beltway problem from combinatorial optimization, and uniform uncertainty principles from harmonic analysis. Our techniques include a certain enhanced form of the probabilistic method, which might be of general interest in its own right.

math.ST

Disordered complex networks: energy optimal lattices and persistent homology

Disordered complex networks are of fundamental interest as stochastic models for information transmission over wireless networks. Well-known networks based on the Poisson point process model have limitations vis-a-vis network efficiency, whereas strongly correlated alternatives, such as those based on random matrix spectra (RMT), have tractability and robustness issues. In this work, we demonstrate that network models based on random perturbations of Euclidean lattices interpolate between Poisson and rigidly structured networks, and allow us to achieve the best of both worlds : significantly improve upon the Poisson model in terms of network efficacy measured by the Signal to Interference plus Noise Ratio (abbrv. SINR) and the related concept of coverage probabilities, at the same time retaining a considerable measure of mathematical and computational simplicity and robustness to erasure and noise. We investigate the optimal choice of the base lattice in this model, connecting it to the celebrated problem optimality of Euclidean lattices with respect to the Epstein Zeta function, which is in turn related to notions of lattice energy. This leads us to the choice of the triangular lattice in 2D and face centered cubic lattice in 3D. We demonstrate that the coverage probability decreases with increasing strength of perturbation, eventually converging to that of the Poisson network. In the regime of low disorder, we approximately characterize the statistical law of the coverage function. In 2D, we determine the disorder strength at which the PTL and the RMT networks are the closest measured by comparing their network topologies via a comparison of their Persistence Diagrams . We demonstrate that the PTL network at this disorder strength can be taken to be an effective substitute for the RMT network model, while at the same time offering the advantages of greater tractability.

eess.SP

Maximum Likelihood under constraints: Degeneracies and Random Critical Points

We investigate the problem of semi-parametric maximum likelihood under constraints on summary statistics. Such a procedure results in a discrete probability distribution that maximises the likelihood among all such distributions under the specified constraints (called estimating equations), and is an approximation to the underlying population distribution. The study of such empirical likelihood originates from the seminal work of Owen. We investigate this procedure in the setting of mis-specified (or biased) estimating equations, i.e. when the null hypothesis is not true. We establish that the behaviour of the optimal distribution under such mis-specification differ markedly from their properties under the null, i.e. when the estimating equations are unbiased and correctly specified. This is manifested by certain degeneracies in the optimal distribution which define the likelihood. Such degeneracies are not observed under the null. Furthermore, we establish an anomalous behaviour of the log-likelihood based Wilks statistic, which, unlike under the null, does not exhibit a chi-squared limit. In the Bayesian setting, we rigorously establish the posterior consistency of procedures based on these ideas, where instead of a parametric likelihood, an empirical likelihood is used to define the posterior distribution. In particular, we show that this posterior, as a random probability measure, rapidly converges to the delta measure at the true parameter value. A novel feature of our approach is the investigation of critical points of random functions in the context of such empirical likelihood. In particular, we obtain the location and the mass of the degenerate optimal weights as the leading and sub-leading terms in a canonical expansion of a particular critical point of a random function that is naturally associated with the model.

math.ST

Point processes, hole events, and large deviations: random complex zeros and Coulomb gases

We consider particle systems (also known as point processes) on the line and in the plane, and are particularly interested in "hole" events, when there are no particles in a large disk (or some other domain). We survey the extensive work on hole probabilities and the related large deviation principles (LDP), which has been undertaken mostly in the last two decades. We mainly focus on the recent applications of LDP-inspired techniques to the study of hole probabilities, and the determination of the most likely configurations of particles that have large holes. As an application of this approach, we illustrate how one can confirm some of the predictions of Jancovici, Lebowitz, and Manificat for large fluctuation in the number of points for the (two-dimensional) $\beta$-Ginibre ensembles. We also discuss some possible directions for future investigations.

math.PR

An easy-to-use empirical likelihood ABC method

Many scientifically well-motivated statistical models in natural, engineering and environmental sciences are specified through a generative process, but in some cases it may not be possible to write down a likelihood for these models analytically. Approximate Bayesian computation (ABC) methods, which allow Bayesian inference in these situations, are typically computationally intensive. Recently, computationally attractive empirical likelihood based ABC methods have been suggested in the literature. These methods heavily rely on the availability of a set of suitable analytically tractable estimating equations. We propose an easy-to-use empirical likelihood ABC method, where the only inputs required are a choice of summary statistic, it's observed value, and the ability to simulate summary statistics for any parameter value under the model. It is shown that the posterior obtained using the proposed method is consistent, and its performance is explored using various examples.

stat.CO

Generalized stealthy hyperuniform processes : maximal rigidity and the bounded holes conjecture

We study translation invariant stochastic processes on $\mathbb{R}^d$ or $\mathbb{Z}^d$ whose diffraction spectrum or structure function $S(k)$, i.e. the Fourier transform of the truncated total pair correlation function, vanishes on an open set $U$ in the wave space. A key family of such processes are stealthy hyperuniform point processes, for which the origin $k=0$ is in $U$; these are of much current physical interest. We show that all such processes exhibit the following remarkable maximal rigidity : namely, the configuration outside a bounded region determines, with probability 1, the exact value (or the exact locations of the points) of the process inside the region. In particular, such processes are completely determined by their tail. In the 1D discrete setting (i.e. $\mathbb{Z}$-valued processes on $\mathbb{Z}$), this can also be seen as a consequence of a recent theorem of Borichev, Sodin and Weiss; in higher dimensions or in the continuum, such a phenomenon seems novel. For stealthy hyperuniform point processes, we prove the Zhang-Stillinger-Torquato conjecture that such processes have bounded holes (empty regions), with a universal bound that depends inversely on the size of $U$.

math.PR

Fluctuations, large deviations and rigidity in hyperuniform systems: a brief survey

We present a brief survey of fluctuations and large deviations of particle systems with subextensive growth of the variance. These are called hyperuniform (or superhomogeneous) systems. We then discuss the relation between hyperuniformity and rigidity. In particular we give sufficient conditions for rigidity of such systems in d=1,2.

math.PR

Number rigidity in superhomogeneous random point fields

We give sufficient conditions for the number rigidity of a translation invariant or periodic point process on $\mathbb{R}^d$, where $d=1,2$. That is, the probability distribution of the number of particles in a bounded domain $\Lambda \subset \mathbb{R}^d$, conditional on the configuration on $\Lambda^\complement$, is concentrated on a single integer $N_\Lambda$. These conditions are : (a) the variance of the number of particles in a bounded domain $\mathcal{O} \subset \mathbb{R}^d$ grows slower than the volume of $\mathcal{O}$ (a.k.a. superhomogeneous point processes), when $\mathcal{O} \uparrow \mathbb{R}^d$ (in a self-similar manner), and (b) the truncated pair correlation function is bounded by $C_1[|x-y|+1]^{-2}$ in $d=1$ and by $C_2[|x-y|+1]^{-(4+\epsilon)}$ in $d=2$. These conditions are satisfied by all known processes with number rigidity ([GP],[G],[PS],[AM],[Bu],[BuDQ], [BBNY], and many more) in $d=1,2$. We also observe, in the light of the results of [PS], that no such criteria exist in $d>2$.

math.PR

Rigidity hierarchy in random point fields: random polynomials and determinantal processes

In certain point processes, the configuration of points outside a bounded domain determines, with probability 1, certain statistical features of the points within the domain. This notion, called rigidity, was introduced in a work of Ghosh and Peres. In this paper, rigidity and the related notion of tolerance are examined systematically and point processes with rigidity of various degrees are introduced. Natural classes of point processes such as determinantal point processes, zero sets of Gaussian entire functions and perturbed lattices are examined from the point of view of rigidity, and general conditions are provided for them to exhibit specified nature of spatially rigid behaviour. In particular, we examine the rigidity of determinantal point processes in terms of their kernel, and demonstrate that a necessary condition for determinantal processes to exhibit rigidity is that their kernel must be a projection. We introduce a one parameter family of point processes which exhibit arbitrarily high levels of rigidity (depending on the choice of parameter value), answering a natural question on point processes with higher levels of rigidity (beyond the known examples of rigidity of local mass and center of mass). Our one parameter family is also related to a natural extension of the standard planar Gaussian analytic function process and their zero sets.

math.PR

Palm measures and rigidity phenomena in point processes

We study the mutual regularity properties of Palm measures of point processes, and establish that a key determining factor for these properties is the rigidity-tolerance behaviour of the point process in question (for those processes that exhibit such behaviour). Thereby, we extend the results of Osada-Shirai, Bufetov and Olshanski to new ensembles, particularly those that are devoid of any determinantal structure. These include the zeroes of the standard planar Gaussian analytic function and several others.

math.PR