Searcharxiv⌕ Search

arXiv subjects

Sumit Mukherjee

Publications and source records attributed to Sumit Mukherjee.

At least 37 records · Page 2Linked to original sources

CLT in high-dimensional Bayesian linear regression with low SNR

We study central limit theorems for linear statistics in high-dimensional Bayesian linear regression with product priors. Unlike the existing literature where the focus is on posterior contraction, we work under a non-contracting regime where neither the likelihood nor the prior dominates the other. This is motivated by modern high-dimensional datasets characterized by a bounded signal-to-noise ratio. This work takes a first step towards understanding limit distributions for one-dimensional projections of the posterior, as well as the posterior mean, in such regimes. Analogous to contractive settings, the resulting limiting distributions are Gaussian, but they heavily depend on the chosen prior and center around the Mean-Field approximation of the posterior. We study two concrete models of interest to illustrate this phenomenon -- the white noise design, and the (misspecified) Bayesian model. As an application, we construct credible intervals and compute their coverage probability under any misspecified prior. Our proofs rely on a combination of recent developments in Berry-Esseen type bounds for Random Field Ising models and both first and second order Poincaré inequalities. Notably, our results do not require any sparsity assumptions on the prior.

math.ST↗

Variational Inference for Latent Variable Models in High Dimensions

Variational inference (VI) is a popular method for approximating intractable posterior distributions in Bayesian inference and probabilistic machine learning. In this paper, we introduce a general framework for quantifying the statistical accuracy of mean-field variational inference (MFVI) for posterior approximation in Bayesian latent variable models with categorical local latent variables (and arbitrary global latent variables). Utilizing our general framework, we capture the exact regime where MFVI 'works' for the celebrated latent Dirichlet allocation model. Focusing on the mixed membership stochastic blockmodel, we show that the vanilla fully factorized MFVI, often used in the literature, is suboptimal. We propose a partially grouped VI algorithm for this model and show that it works, and derive its exact finite-sample performance. We further illustrate that our bounds are tight for both the above models. Our proof techniques, which extend the framework of nonlinear large deviations, open the door for the analysis of MFVI in other latent variable models.

math.ST↗

Large random matrices with given margins

We study large random matrices with i.i.d. entries conditioned to have prescribed row and column sums (margins), a problem connected to relative entropy minimization, Schrödinger bridges, contingency tables, and random graphs with given degree sequences. Our central result is a `transference principle': the complex margin-conditioned matrix can be closely approximated by a simpler matrix whose entries are independent and drawn from an exponential tilting of the original model. The tilt parameters are determined by the sum of two potentials. We establish phase diagrams for `tame margins', where these potentials are uniformly bounded. This framework resolves a 2011 conjecture by Chatterjee, Diaconis, and Sly on $δ$-tame degree sequences and generalizes a sharp phase transition in contingency tables obtained by Dittmer, Lyu, and Pak in 2020. For tame margins, we show that a generalized Sinkhorn algorithm can compute the potentials at a dimension-free exponential rate. Our limit theory further establishes that for a convergent sequence of tame margins, the potentials converge as fast as the margins converge. We apply this framework and obtain several key results for the conditioned matrix: The marginal distribution of any single entry is asymptotically an exponential tilting of the base measure, resolving a 2010 conjecture by Barvinok on contingency tables. The conditioned matrix concentrates in cut norm around a `typical table' (the expectation of the tilted model), which acts as a static Schrödinger bridge between the margins. The empirical singular value distribution of the rescaled matrix converges to an explicit law determined by the variance profile of the tilted model. In particular, we confirm the universality of the Marchenko-Pastur law for constant linear margins.

math.PR↗

Contextuality sans incompatibility in the simplest scenario: Communication supremacy of a qubit

Conventional wisdom asserts that measurement incompatibility is necessary for revealing the non-locality and contextuality. In contrast, a recent work [Phys. Rev. Lett. 130, 230201 (2023)] demonstrates the generalized contextuality without measurement incompatibility by using a five-outcome qubit measurement. In this paper, we introduce a two-party prepare-measure communication game involving specific constraints on preparations, and we demonstrate contextuality sans incompatibility in the simplest measurement scenario, requiring only a three-outcome extremal qubit measurement. This contrasts with the aforementioned five-outcome qubit measurement, which can be simulated by an appropriate convex mixture of five three-outcome incompatible qubit measurements. Furthermore, we illustrate that our result has a prominent implication in information theory. Our communication game can be perceived as a constrained Holevo-Frankle-Weiner (HFW) scenario, as operational restrictions are imposed on preparations. We show that the maximum success probability of the game by using a qubit surpasses that attainable by a c-bit, even when shared randomness is a free resource. Consequently, this finding exemplifies the supremacy of a qubit over a c-bit within a constrained HFW framework. Thus, alongside offering fresh insights into quantum foundations, our results pave a novel pathway for exploring the efficacy of a qubit in information processing tasks.

quant-ph↗

Persistence and Ball Exponents for Gaussian Stationary Processes

Consider a real Gaussian stationary process $f_ρ$, indexed on either $\mathbb{R}$ or $\mathbb{Z}$ and admitting a spectral measure $ρ$. We study $θ_ρ^\ell=-\lim\limits_{T\to\infty}\frac{1}{T} \log\mathbb{P}\left(\inf_{t\in[0,T]}f_ρ(t)>\ell\right)$, the persistence exponent of $f_ρ$. We show that, if $ρ$ has a positive density at the origin, then the persistence exponent exists; moreover, if $ρ$ has an absolutely continuous component, then $θ_ρ^\ell>0$ if and only if this spectral density at the origin is finite. We further establish continuity of $θ_ρ^\ell$ in $\ell$, in $ρ$ (under a suitable metric) and, if $ρ$ is compactly supported, also in dense sampling. Analogous continuity properties are shown for $ψ_ρ^\ell=-\lim\limits_{T\to\infty}\frac{1}{T} \log\mathbb{P}\left(\inf_{t\in[0,T]}|f_ρ(t)|\le \ell\right)$, the ball exponent of $f_ρ$, and it is shown to be positive if and only if $ρ$ has an absolutely continuous component.

math.PR↗

Motif Estimation via Subgraph Sampling: The Fourth Moment Phenomenon

Network sampling is an indispensable tool for understanding features of large complex networks where it is practically impossible to search over the entire graph. In this paper, we develop a framework for statistical inference for counting network motifs, such as edges, triangles, and wedges, in the widely used subgraph sampling model, where each vertex is sampled independently, and the subgraph induced by the sampled vertices is observed. We derive necessary and sufficient conditions for the consistency and the asymptotic normality of the natural Horvitz-Thompson (HT) estimator, which can be used for constructing confidence intervals and hypothesis testing for the motif counts based on the sampled graph. In particular, we show that the asymptotic normality of the HT estimator exhibits an interesting fourth-moment phenomenon, which asserts that the HT estimator (appropriately centered and rescaled) converges in distribution to the standard normal whenever its fourth-moment converges to 3 (the fourth-moment of the standard normal distribution). As a consequence, we derive the exact thresholds for consistency and asymptotic normality of the HT estimator in various natural graph ensembles, such as sparse graphs with bounded degree, Erdos-Renyi random graphs, random regular graphs, and dense graphons.

math.ST↗

Constrained Measurement Incompatibility from Generalised Contextuality of Steered Preparation

In a bipartite Bell scenario involving two local measurements per party and two outcome per measurement, the measurement incompatibility in one wing is both necessary and sufficient to reveal the nonlocality. However, such a one-to-one correspondence fails when one of the observers performs more than two measurements. In such a scenario, the measurement incompatibility is necessary but not sufficient to reveal the nonlocality. In this work, within the formalism of general probabilistic theory (GPT), we demonstrate that unlike the nonlocality, the incompatibility of N arbitrary measurements in one wing is both necessary and sufficient for revealing the generalised contextuality for the sub-system in the other wing. Further, we formulate a novel form of inequality for any GPT that are necessary for N-wise compatibility of N arbitrary observables. Moreover, we argue that any theory that violates the proposed inequality possess a degree of incompatibility that can be quantified through the amount of violation. Finally, we claim that it is the generalised contextuality that provides a restriction to the allowed degree of measurement incompatibility of any viable theory of nature and thereby super-select the the quantum theory.

quant-ph↗

Universality of Persistence of Random Polynomials

We investigate the probability that a random polynomial with independent, mean-zero and finite variance coefficients has no real zeros. Specifically, we consider a random polynomial of degree $2n$ with coefficients given by an i.i.d. sequence of mean-zero, variance-1 random variables, multiplied by an $\fracα{2}$-regularly varying sequence for $α>-1$. We show that the probability of no real zeros is asymptotically $n^{-2(b_α+b_0)}$, where $b_α$ is the persistence exponents of a mean-zero, one-dimensional stationary Gaussian processes with covariance function as $\mathrm{sech}((t-s)/2)^{α+1}$. Our work generalizes the previous results of Dembo et al. [DPSZ02] and Dembo \& Mukherjee [DM15] by removing the requirement of finite moments of all order or Gaussianity. In particular, in the special case $α= 0$, our findings confirm a conjecture by Poonen and Stoll [PS99, Section 9.1] concerning random polynomials with i.i.d. coefficients.

math.PR↗

Synthesizing Proton-Density Fat Fraction and $R_2^*$ from 2-point Dixon MRI with Generative Machine Learning

Magnetic Resonance Imaging (MRI) is the gold standard for measuring fat and iron content non-invasively in the body via measures known as Proton Density Fat Fraction (PDFF) and $R_2^*$, respectively. However, conventional PDFF and $R_2^*$ quantification methods operate on MR images voxel-wise and require at least three measurements to estimate three quantities: water, fat, and $R_2^*$. Alternatively, the two-point Dixon MRI protocol is widely used and fast because it acquires only two measurements; however, these cannot be used to estimate three quantities voxel-wise. Leveraging the fact that neighboring voxels have similar values, we propose using a generative machine learning approach to learn PDFF and $R_2^*$ from Dixon MRI. We use paired Dixon-IDEAL data from UK Biobank in the liver and a Pix2Pix conditional GAN to demonstrate the first large-scale $R_2^*$ imputation from two-point Dixon MRIs. Using our proposed approach, we synthesize PDFF and $R_2^*$ maps that show significantly greater correlation with ground-truth than conventional voxel-wise baselines.

cs.CV↗

Fluctuations of Quadratic Chaos

In this paper we characterize all distributional limits of the random quadratic form $T_n =\sum_{1\le u< v\le n} a_{u, v} X_u X_v$, where $((a_{u, v}))_{1\le u,v\le n}$ is a $\{0, 1\}$-valued symmetric matrix with zeros on the diagonal and $X_1, X_2, \ldots, X_n$ are i.i.d.~ mean $0$ variance $1$ random variables with common distribution function $F$. In particular, we show that any distributional limit of $S_n:=T_n/\sqrt{\mathrm{Var}[T_n]}$ can be expressed as the sum of three independent components: a Gaussian, a (possibly) infinite weighted sum of independent centered chi-squares, and a Gaussian mixture with a random variance. As a consequence, we prove a fourth moment theorem for the asymptotic normality of $S_n$, which applies even when $F$ does not have finite fourth moment. More formally, we show that $S_n$ converges to $N(0, 1)$ if and only if the fourth moment of $S_n$ (appropriately truncated when $F$ does not have finite fourth moment) converges to 3 (the fourth moment of the standard normal distribution).

math.PR↗

On Naive Mean-Field Approximation for high-dimensional canonical GLMs

We study the validity of the Naive Mean Field (NMF) approximation for canonical GLMs with product priors. This setting is challenging due to the non-conjugacy of the likelihood and the prior. Using the theory of non-linear large deviations (Austin 2019, Chatterjee, Dembo 2016, Eldan 2018), we derive sufficient conditions for the tightness of the NMF approximation to the log-normalizing constant of the posterior distribution. As a second contribution, we establish that under minor conditions on the design, any NMF optimizer is a product distribution where each component is a quadratic tilt of the prior. In turn, this suggests novel iterative algorithms for fitting the NMF optimizer to the target posterior. Finally, we establish that if the NMF optimization problem has a "well-separated maximizer", then this optimizer governs the probabilistic properties of the posterior. Specifically, we derive credible intervals with average coverage guarantees, and characterize the prediction performance on an out-of-sample datapoint in terms of this dominant optimizer.

math.ST↗

Device-independent certification of degeneracy-breaking measurements

In a device-independent Bell test, the devices are considered to be black boxes and the dimension of the system remains unspecified. The dichotomic observables involved in such a Bell test can be degenerate and one may invoke a suitable measurement scheme to lift the degeneracy. However, the standard Bell test cannot account for whether or up to what extent the degeneracy is lifted, as the effect of lifting the degeneracy can only be reflected in the post-measurement states, which the standard Bell tests do not certify. In this work, we demonstrate the device-independent certification of degeneracy-breaking measurement based on the sequential Bell test by multiple observers who perform degeneracy-breaking unsharp measurements characterized by positive-operator-valued measures (POVMs) - the noisy variants of projectors. The optimal quantum violation of Clauser-Horne-Shimony-Holt inequality by multiple sequential observers eventually enables us to certify up to what extent the degeneracy has been lifted. In particular, our protocol certifies the upper bound on the number of POVMs used for performing such measurements along with the entangled state and measurement observables. We use an elegant sum-of-squares approach that powers such certification of degeneracy-breaking measurements.

quant-ph↗

Assessment of Differentially Private Synthetic Data for Utility and Fairness in End-to-End Machine Learning Pipelines for Tabular Data

Differentially private (DP) synthetic data sets are a solution for sharing data while preserving the privacy of individual data providers. Understanding the effects of utilizing DP synthetic data in end-to-end machine learning pipelines impacts areas such as health care and humanitarian action, where data is scarce and regulated by restrictive privacy laws. In this work, we investigate the extent to which synthetic data can replace real, tabular data in machine learning pipelines and identify the most effective synthetic data generation techniques for training and evaluating machine learning models. We investigate the impacts of differentially private synthetic data on downstream classification tasks from the point of view of utility as well as fairness. Our analysis is comprehensive and includes representatives of the two main types of synthetic data generation algorithms: marginal-based and GAN-based. To the best of our knowledge, our work is the first that: (i) proposes a training and evaluation framework that does not assume that real data is available for testing the utility and fairness of machine learning models trained on synthetic data; (ii) presents the most extensive analysis of synthetic data set generation algorithms in terms of utility and fairness when used for training machine learning models; and (iii) encompasses several different definitions of fairness. Our findings demonstrate that marginal-based synthetic data generators surpass GAN-based ones regarding model training utility for tabular data. Indeed, we show that models trained using data generated by marginal-based algorithms can exhibit similar utility to models trained using real data. Our analysis also reveals that the marginal-based synthetic data generator MWEM PGM can train models that simultaneously achieve utility and fairness characteristics close to those obtained by models trained with real data.

cs.LG↗

Tensile quantum-to-classical transition of macroscopic entangled states under complete coarse-grained measurements

The macroscopic limit at which the quantum-to-classical transition occurs remains as one of the long-standing questions in the foundations of quantum theory. There are evidences that the macroscopic limit to which the quantumness of a system persists depends on the degree of interaction due to the measurement processes. For instance, with a system having a considerably large Hilbert space dimension, if the measurement is performed in such a way that the outcome of the measurement only reveals a coarse-grained version of the information about the individual level of the concerned system then the disturbance due to the measurement process can be considered to be infinitesimally small. Based on such coarse-grained measurement the dependence of Bell inequality violation on the degree of coarsening has already been investigated [Phys. Rev. Lett. 112, 010402 (2014)]. In this paper, we first capture the fact that when local-realism is taken to be the defining notion of classicality, the effect of the degree of coarsening on the downfall of quantumness of a macroscopic entangled state can be compensated by testing a Bell-inequality of a higher number of settings from a family of symmetric Bell-inequalities if the number of settings is odd. However, on the contrary, we show that such compensation can not be seen when we witness such quantum-to-classical transition using symmetric Bell inequalities having an even number of settings. Finally, complementing the above result, we show that when unsteerability is taken as the classicality, for both odd and even numbers of settings the degree of coarsening at which the quantum-to-classical transition occurs can be consistently pushed ahead by testing a linear steering inequality of a higher number of settings and observing its violation. We further extend our treatment for mixed macroscopic entangled states

quant-ph↗

Large scale synthesis of 2D graphene oxide by mechanical milling of 3D carbon nanoparticles in air

Graphene oxide (GO) is one of the important functional materials. Large-scale synthesis of it is very challenging. Following a simple cost-effective route, large-scale GO was produced by mechanical (ball) milling, in air, of carbon nanoparticles (CNPs) present in carbon soot in the present study. The thickness of the GO layer was seen to decrease with an increase in milling time. Ball milling provided the required energy to acquire the in-plane graphitic order in the CNPs reducing the disorders in it. As the surface area of the layered structure became more and more with the increase in milling time, more and more oxygen of air got attached to the carbon in graphene leading to the formation of GO. An increase in the time of the ball mill up to 5 hours leads to a significant increase in the content of GO. Thus ball milling can be useful to produce large-scale two-dimensional GO for a short time.

cond-mat.mtrl-sci↗

Ising Models on Dense Regular Graphs

In this paper, we derive the limit of experiments for one parameter Ising models on dense regular graphs. In particular, we show that the limiting experiment is Gaussian in the low temperature regime, non Gaussian in the critical regime, and an infinite collection of Gaussians in the high temperature regime. We also derive the limiting distributions of the maximum likelihood and maximum pseudo-likelihood estimators, and study limiting power for tests of hypothesis against contiguous alternatives (whose scaling changes across the regimes). To the best of our knowledge, this is the first attempt at establishing the classical limits of experiments for Ising models (and more generally, Markov random fields).

math.ST↗

Inference on a class of exponential families on permutations

In this paper we study a class of exponential family on permutations, which includes some of the commonly studied Mallows models. We show that the pseudo-likelihood estimator for the natural parameter in the exponential family is asymptotically normal, with an explicit variance. Using this, we are able to construct asymptotically valid confidence intervals. We also show that the MLE for the same problem is consistent everywhere, and asymptotically normal at the origin. In this special case, the asymptotic variance of the cost effective pseudo-likelihood estimator turns out to be the same as the cost prohibitive MLE. To the best of our knowledge, this is the first inference result on permutation models including Mallows models, excluding the very special case of Mallows model with Kendall's Tau.

math.ST↗

Large deviation principle for random permutations

We derive a large deviation principle for random permutations induced by probability measures of the unit square, called permutons. These permutations are called $μ$-random permutations. We also introduce and study a new general class of models of random permutations, called Gibbs permutation models, which combines and generalizes $μ$-random permutations and the celebrated Mallows model for permutations. Most of our results hold in the general setting of Gibbs permutation models. We apply the tools that we develop to the case of $μ$-random permutations conditioned to have an atypical proportion of patterns. Several results are made more concrete in the specific case of inversions. For instance, we prove the existence of at least one phase transition for a generalized version of the Mallows model where the base measure is non-uniform. This is in contrast with the results of Starr (2009, 2018) on the (standard) Mallows model, where the absence of phase transition, i.e., phase uniqueness, was proven. Our results naturally lead us to investigate a new notion of permutons, called conditionally constant permutons, which generalizes both pattern-avoiding and pattern-packing permutons. We describe some properties of conditionally constant permutons with respect to inversions. The study of conditionally constant permutons for general patterns seems to be a challenging problem.

math.PR↗