Searcharxiv⌕ Search

arXiv subjects

Sumit Mukherjee

Publications and source records attributed to Sumit Mukherjee.

At least 55 records · Page 3Linked to original sources

Persistence probabilities of weighted sums of stationary Gaussian sequences

With $\{ξ_i\}_{i\ge 0}$ being a centered stationary Gaussian sequence with non-negative correlation function $ρ(i):=\mathbb{E}[ ξ_0ξ_i]$ and $\{σ(i)\}_{i\ge 1}$ a sequence of positive reals, we study the asymptotics of the persistence probability of the weighted sum $\sum_{i=1}^\ell σ(i) ξ_i$, $\ell\ge 1$. For summable correlations $ρ$, we show that the persistence exponent is universal. On the contrary, for non-summable $ρ$, even for polynomial weight functions $σ(i)\sim i^p$ the persistence exponent depends on the rate of decay of the correlations (encoded by a parameter $H$) and on the polynomial rate $p$ of $σ$. In this case, we show existence of the persistence exponent $θ(H,p)$ and study its properties as a function of $(p,H)$. During the course of our proofs, we develop several tools for dealing with exit problems for Gaussian processes with non-negative correlations -- e.g.\ a continuity result for persistence exponents and a necessary and sufficient criterion for the persistence exponent to be zero -- that might be of independent interest.

math.PR↗

Robust certification of unsharp instruments through sequential quantum advantages in a prepare-measure communication game

Communication games are one of the widely used tools that are designed to demonstrate quantum supremacy over classical resources. In that, two or more parties collaborate to perform an information processing task to achieve the highest success probability of winning the game. We propose a specific two-party communication game in the prepare-measure scenario that relies on an encoding-decoding task of specific information. We first demonstrate that quantum theory outperforms the classical preparation non-contextual theory, and the optimal quantum success probability of such a communication game enables the semi-device-independent certification of qubit states and measurements. Further, we consider the sequential sharing of quantum preparation contextuality and show that, at most, two sequential observers can share the quantum advantage. The sub-optimal quantum advantages for two sequential observers form an optimal pair that certifies a unique value of the unsharpness parameter of the first observer. Since the practical implementation inevitably introduces noise, we devised a scheme to demonstrate the robust certification of the states and unsharp measurement instruments of both the sequential observers.

quant-ph↗

MACE: A Flexible Framework for Membership Privacy Estimation in Generative Models

Generative machine learning models are being increasingly viewed as a way to share sensitive data between institutions. While there has been work on developing differentially private generative modeling approaches, these approaches generally lead to sub-par sample quality, limiting their use in real world applications. Another line of work has focused on developing generative models which lead to higher quality samples but currently lack any formal privacy guarantees. In this work, we propose the first formal framework for membership privacy estimation in generative models. We formulate the membership privacy risk as a statistical divergence between training samples and hold-out samples, and propose sample-based methods to estimate this divergence. Compared to previous works, our framework makes more realistic and flexible assumptions. First, we offer a generalizable metric as an alternative to the accuracy metric especially for imbalanced datasets. Second, we loosen the assumption of having full access to the underlying distribution from previous studies , and propose sample-based estimations with theoretical guarantees. Third, along with the population-level membership privacy risk estimation via the optimal membership advantage, we offer the individual-level estimation via the individual privacy risk. Fourth, our framework allows adversaries to access the trained model via a customized query, while prior works require specific attributes.

cs.CR↗

Fluctuations in Mean-Field Ising models

In this paper, we study the fluctuations of the average magnetization in an Ising model on an approximately $d_N$ regular graph $G_N$ on $N$ vertices. In particular, if $G_N$ is \enquote{well connected}, we show that whenever $d_N\gg \sqrt{N}$, the fluctuations are universal and same as that of the Curie-Weiss model in the entire Ferro-magnetic parameter regime. We give a counterexample to demonstrate that the condition $d_N\gg \sqrt{N}$ is tight, in the sense that the limiting distribution changes if $d_N\sim \sqrt{N}$ except in the high temperature regime. By refining our argument, we extend universality in the high temperature regime up to $d_N\gg N^{1/3}$. Our results conclude universal fluctuations of the average magnetization in Ising models on regular graphs, Erdős-Rényi graphs (directed and undirected), stochastic block models, and sparse regular graphons. In fact, our results apply to general matrices with non-negative entries, including Ising models on a Wigner matrix, and the block spin Ising model. As a by-product of our proof technique, we obtain Berry-Esseen bounds for these fluctuations, exponential concentration for the average of spins, and tight error bounds for the Mean-Field approximation of the partition function.

math.PR↗

Mean field approximations via log-concavity

We propose a new approach to deriving quantitative mean field approximations for any probability measure $P$ on $\mathbb{R}^n$ with density proportional to $e^{f(x)}$, for $f$ strongly concave. We bound the mean field approximation for the log partition function $\log \int e^{f(x)}dx$ in terms of $\sum_{i \neq j}\mathbb{E}_{Q^*}|\partial_{ij}f|^2$, for a semi-explicit probability measure $Q^*$ characterized as the unique mean field optimizer, or equivalently as the minimizer of the relative entropy $H(\cdot\,|\,P)$ over product measures. This notably does not involve metric-entropy or gradient-complexity concepts which are common in prior work on nonlinear large deviations. Three implications are discussed, in the contexts of continuous Gibbs measures on large graphs, high-dimensional Bayesian linear regression, and the construction of decentralized near-optimizers in high-dimensional stochastic control problems. Our arguments are based primarily on functional inequalities and the notion of displacement convexity from optimal transport.

math.PR↗

Semi-device independent certification of multiple unsharpness parameters through sequential measurements

Based on a sequential communication game, semi-device independent certification of an unsharp instrument has recently been demonstrated [\href{https://iopscience.iop.org/article/10.1088/1367-2630/ab3773}{New J. Phys. 21 083034 (2019), }\href{https://journals.aps.org/prresearch/abstract/10.1103/PhysRevResearch.2.033014}{ Phys. Rev. Research 2, 033014 (2020)}]. In this paper, we provide semi-device independent self-testing protocols in the prepare-measure scenario to certify multiple unsharpness parameters along with the states and the measurement settings. This is achieved through the sequential quantum advantage shared by multiple independent observers in a suitable communication game known as parity-oblivious random-access-code. We demonstrate that in 3-bit parity-oblivious random-access-code, at most three independent observers can sequentially share quantum advantage. The optimal pair (triple) of quantum advantages enables us to uniquely certify the qubit states, the measurement settings, and the unsharpness parameter(s). The practical implementation of a given protocol involves inevitable losses. In a sub-optimal scenario, we derive a certified interval within which a specific unsharpness parameter has to be confined. We extend our treatment to the 4-bit case and show that at most two observers can share quantum advantage for the qubit system. Further, we provide a sketch to argue that four sequential observers can share the quantum advantage for the two-qubit system, thereby enabling the certification of three unsharpness parameters.

quant-ph↗

Discriminating mirror symmetric states with restricted contextual advantage

The generalized notion of noncontextuality provides an avenue to explore the fundamental departure of quantum theory from a classical explanation. Recently, extracting a different form of quantum advantage in various information processing tasks has received an upsurge of interest. In a recent work [D. Schmid and R. W. Spekkens, Phys. Rev. X 8, 011015 (2018)] it has been demonstrated that discrimination of two nonorthogonal pure quantum states entails contextual advantage when the states are supplied with equal prior probabilities. We generalized the work to arbitrary prior probabilities as well as to three arbitrary mirror-symmetric states. We show that the contextual advantage can be obtained for any value of prior probability when only two quantum states are present in the task. But surprisingly, in the case of three mirror-symmetric states, the contextual advantage is available only for a restrictive range of prior probabilities with which the states are prepared.

quant-ph↗

Reducing bias and increasing utility by federated generative modeling of medical images using a centralized adversary

We introduce FELICIA (FEderated LearnIng with a CentralIzed Adversary) a generative mechanism enabling collaborative learning. In particular, we show how a data owner with limited and biased data could benefit from other data owners while keeping data from all the sources private. This is a common scenario in medical image analysis where privacy legislation prevents data from being shared outside local premises. FELICIA works for a large family of Generative Adversarial Networks (GAN) architectures including vanilla and conditional GANs as demonstrated in this work. We show that by using the FELICIA mechanism, a data owner with limited image samples can generate high-quality synthetic images with high utility while neither data owners has to provide access to its data. The sharing happens solely through a central discriminator that has access limited to synthetic data. Here, utility is defined as classification performance on a real test set. We demonstrate these benefits on several realistic healthcare scenarions using benchmark image datasets (MNIST, CIFAR-10) as well as on medical images for the task of skin lesion classification. With multiple experiments, we show that even in the worst cases, combining FELICIA with real data gracefully achieves performance on par with real data while most results significantly improves the utility.

stat.ML↗

Signal Detection in Degree Corrected ERGMs

In this paper, we study sparse signal detection problems in Degree Corrected Exponential Random Graph Models (ERGMs). We study the performance of two tests based on the conditionally centered sum of degrees and conditionally centered maximum of degrees, for a wide class of such ERGMs. The performance of these tests match the performance of the corresponding uncentered tests in the $β$ model. Focusing on the degree corrected two star ERGM, we show that improved detection is possible at criticality using a test based on (unconditional) sum of degrees. In this setting we provide matching lower bounds in all parameter regimes, which is based on correlations estimates between degrees under the alternative, and of possible independent interest.

math.ST↗

An Analysis of the Deployment of Models Trained on Private Tabular Synthetic Data: Unexpected Surprises

Diferentially private (DP) synthetic datasets are a powerful approach for training machine learning models while respecting the privacy of individual data providers. The effect of DP on the fairness of the resulting trained models is not yet well understood. In this contribution, we systematically study the effects of differentially private synthetic data generation on classification. We analyze disparities in model utility and bias caused by the synthetic dataset, measured through algorithmic fairness metrics. Our first set of results show that although there seems to be a clear negative correlation between privacy and utility (the more private, the less accurate) across all data synthesizers we evaluated, more privacy does not necessarily imply more bias. Additionally, we assess the effects of utilizing synthetic datasets for model training and model evaluation. We show that results obtained on synthetic data can misestimate the actual model performance when it is deployed on real data. We hence advocate on the need for defining proper testing protocols in scenarios where differentially private synthetic datasets are utilized for model training and evaluation.

stat.ML↗

A machine learning pipeline for aiding school identification from child trafficking images

Child trafficking in a serious problem around the world. Every year there are more than 4 million victims of child trafficking around the world, many of them for the purposes of child sexual exploitation. In collaboration with UK Police and a non-profit focused on child abuse prevention, Global Emancipation Network, we developed a proof-of-concept machine learning pipeline to aid the identification of children from intercepted images. In this work, we focus on images that contain children wearing school uniforms to identify the school of origin. In the absence of a machine learning pipeline, this hugely time consuming and labor intensive task is manually conducted by law enforcement personnel. Thus, by automating aspects of the school identification process, we hope to significantly impact the speed of this portion of child identification. Our proposed pipeline consists of two machine learning models: i) to identify whether an image of a child contains a school uniform in it, and ii) identification of attributes of different school uniform items (such as color/texture of shirts, sweaters, blazers etc.). We describe the data collection, labeling, model development and validation process, along with strategies for efficient searching of schools using the model predictions.

cs.CV↗

Becoming Good at AI for Good

AI for good (AI4G) projects involve developing and applying artificial intelligence (AI) based solutions to further goals in areas such as sustainability, health, humanitarian aid, and social justice. Developing and deploying such solutions must be done in collaboration with partners who are experts in the domain in question and who already have experience in making progress towards such goals. Based on our experiences, we detail the different aspects of this type of collaboration broken down into four high-level categories: communication, data, modeling, and impact, and distill eleven takeaways to guide such projects in the future. We briefly describe two case studies to illustrate how some of these takeaways were applied in practice during our past collaborations.

cs.CY↗

Variational Inference in high-dimensional linear regression

We study high-dimensional Bayesian linear regression with product priors. Using the nascent theory of non-linear large deviations (Chatterjee and Dembo,2016), we derive sufficient conditions for the leading-order correctness of the naive mean-field approximation to the log-normalizing constant of the posterior distribution. Subsequently, assuming a true linear model for the observed data, we derive a limiting infinite dimensional variational formula for the log normalizing constant of the posterior. Furthermore, we establish that under an additional "separation" condition, the variational problem has a unique optimizer, and this optimizer governs the probabilistic properties of the posterior distribution. We provide intuitive sufficient conditions for the validity of this "separation" condition. Finally, we illustrate our results on concrete examples with specific design matrices.

math.ST↗

Statistics of the two-star ERGM

In this paper, we explore the two-star Exponential Random Graph Model, which is a two parameter exponential family on the space of simple labeled graphs. We introduce auxiliary variables to express the two-star model as a mixture of the $β$ model on networks. Using this representation, we study asymptotic distribution of the number of edges, and the sampling variance of the degrees. In particular, the limiting distribution for the number of edges has similar phase transition behavior to that of the magnetization in the Curie-Weiss Ising model of Statistical Physics. Using this, we show existence of consistent estimates for both parameters in all parameter domains. Finally, we prove that the centered partial sum of degrees converges as a process to a Brownian bridge in all parameter domains, irrespective of the phase transition.

math.ST↗

Monochromatic Subgraphs in Randomly Colored Graphons

Let $T(H, G_n)$ be the number of monochromatic copies of a fixed connected graph $H$ in a uniformly random coloring of the vertices of the graph $G_n$. In this paper we give a complete characterization of the limiting distribution of $T(H, G_n)$, when $\{G_n\}_{n \geq 1}$ is a converging sequence of dense graphs. When the number of colors grows to infinity, depending on whether the expected value remains bounded, $T(H, G_n)$ either converges to a finite linear combination of independent Poisson variables or a normal distribution. On the other hand, when the number of colors is fixed, $T(H, G_n)$ converges to a (possibly infinite) linear combination of independent centered chi-squared random variables. This generalizes the classical birthday problem, which involves understanding the asymptotics of $T(K_s, K_n)$, the number of monochromatic $s$-cliques in a complete graph $K_n$ ($s$-matching birthdays among a group of $n$ friends), to general monochromatic subgraphs in a network.

math.PR↗

privGAN: Protecting GANs from membership inference attacks at low cost

Generative Adversarial Networks (GANs) have made releasing of synthetic images a viable approach to share data without releasing the original dataset. It has been shown that such synthetic data can be used for a variety of downstream tasks such as training classifiers that would otherwise require the original dataset to be shared. However, recent work has shown that the GAN models and their synthetically generated data can be used to infer the training set membership by an adversary who has access to the entire dataset and some auxiliary information. Current approaches to mitigate this problem (such as DPGAN) lead to dramatically poorer generated sample quality than the original non--private GANs. Here we develop a new GAN architecture (privGAN), where the generator is trained not only to cheat the discriminator but also to defend membership inference attacks. The new mechanism provides protection against this mode of attack while leading to negligible loss in downstream performances. In addition, our algorithm has been shown to explicitly prevent overfitting to the training set, which explains why our protection is so effective. The main contributions of this paper are: i) we propose a novel GAN architecture that can generate synthetic data in a privacy preserving manner without additional hyperparameter tuning and architecture selection, ii) we provide a theoretical understanding of the optimal solution of the privGAN loss function, iii) we demonstrate the effectiveness of our model against several white and black--box attacks on several benchmark datasets, iv) we demonstrate on three common benchmark datasets that synthetic images generated by privGAN lead to negligible loss in downstream performance when compared against non--private GANs.

cs.LG↗

Detecting Structured Signals in Ising Models

In this paper, we study the effect of dependence on detecting a class of signals in Ising models, where the signals are present in a structured way. Examples include Ising Models on lattices, and Mean-Field type Ising Models (Erdős-Rényi, Random regular, and dense graphs). Our results rely on correlation decay and mixing type behavior for Ising Models, and demonstrate the beneficial behavior of criticality in the detection of strictly lower signals. As a by-product of our proof technique, we develop sharp control on mixing and spin-spin correlation for several Mean-Field type Ising Models in all regimes of temperature -- which might be of independent interest.

math.PR↗

Persistence exponents in Markov chains

We prove the existence of the persistence exponent $$\logλ:=\lim_{n\to\infty}\frac{1}{n}\log \mathbb{P}_μ(X_0\in S,\ldots,X_n\in S)$$ for a class of time homogeneous Markov chains $\{X_i\}_{i\geq 0}$ taking values in a Polish space, where $S$ is a Borel measurable set and $μ$ is an initial distribution. Focusing on the case of AR($p$) and MA($q$) processes with $p,q\in \mathbb{N}$ and continuous innovation distribution, we study the existence of $λ$ and its continuity in the parameters of the AR and MA processes, respectively, for $S=\mathbb{R}_{\geq 0}$. For AR processes with log-concave innovation distribution, we prove the strict monotonicity of $λ$. Finally, we compute new explicit exponents in several concrete examples.

math.PR↗