Searcharxiv⌕ Search

arXiv subjects

Sayan Mukherjee

Publications and source records attributed to Sayan Mukherjee.

At least 37 records · Page 2Linked to original sources

Irreducibility of Markov Chains on simplicial complexes, the Spectrum of the Discrete Hodge Laplacian and Homology

Random walks on graphs are a fundamental concept in graph theory and play a crucial role in solving a wide range of theoretical and applied problems in discrete math, probability, theoretical computer science, network science, and machine learning. The connection between Markov chains on graphs and their geometric and topological structures is the main reason why such a wide range of theoretical and practical applications exist. Graph connectedness ensures irreducibility of a Markov chain. The convergence rate to the stationary distribution is determined by the spectrum of the graph Laplacian which is associated with lower bounds on graph curvature. Furthermore, walks on graphs are used to infer structural properties of underlying manifolds in data analysis and manifold learning. However, an important question remains: can similar connections be established between Markov chains on simplicial complexes and the topology, geometry, and spectral properties of complexes? Additionally, can we gain topological, geometric, or analytic information about a manifold by defining appropriate Markov chains on its triangulations? These questions are not only theoretically important but answers to them provide powerful tools for the analysis of complex networks that go beyond the analysis of pairwise interactions. In this paper, we provide an integrated overview of the existing results on random walks on simplicial complexes, using the novel perspective of signed graphs. This perspective sheds light on previously unknown aspects such as irreducibility conditions. We show that while up-walks on higher dimensional simplexes can never be irreducible, the down walks become irreducible if and only if the complex is orientable. We believe that this new integrated perspective can be extended beyond discrete structures and enables exploration of classical problems for triangulable manifolds.

math.SP↗

Asymptotics of Bayesian Uncertainty Estimation in Random Features Regression

In this paper we compare and contrast the behavior of the posterior predictive distribution to the risk of the maximum a posteriori estimator for the random features regression model in the overparameterized regime. We will focus on the variance of the posterior predictive distribution (Bayesian model average) and compare its asymptotics to that of the risk of the MAP estimator. In the regime where the model dimensions grow faster than any constant multiple of the number of samples, asymptotic agreement between these two quantities is governed by the phase transition in the signal-to-noise ratio. They also asymptotically agree with each other when the number of samples grow faster than any constant multiple of model dimensions. Numerical simulations illustrate finer distributional properties of the two quantities for finite dimensions. We conjecture they have Gaussian fluctuations and exhibit similar properties as found by previous authors in a Gaussian sequence model, which is of independent theoretical interest.

stat.ML↗

Generalized Bayes Approach to Inverse Problems with Model Misspecification

We propose a general framework for obtaining probabilistic solutions to PDE-based inverse problems. Bayesian methods are attractive for uncertainty quantification but assume knowledge of the likelihood model or data generation process. This assumption is difficult to justify in many inverse problems, where the specification of the data generation process is not obvious. We adopt a Gibbs posterior framework that directly posits a regularized variational problem on the space of probability distributions of the parameter. We propose a novel model comparison framework that evaluates the optimality of a given loss based on its "predictive performance". We provide cross-validation procedures to calibrate the regularization parameter of the variational objective and compare multiple loss functions. Some novel theoretical properties of Gibbs posteriors are also presented. We illustrate the utility of our framework via a simulated example, motivated by dispersion-based wave models used to characterize arterial vessels in ultrasound vibrometry.

stat.ME↗

Bi-Objective Lexicographic Optimization in Markov Decision Processes with Related Objectives

We consider lexicographic bi-objective problems on Markov Decision Processes (MDPs), where we optimize one objective while guaranteeing optimality of another. We propose a two-stage technique for solving such problems when the objectives are related (in a way that we formalize). We instantiate our technique for two natural pairs of objectives: minimizing the (conditional) expected number of steps to a target while guaranteeing the optimal probability of reaching it; and maximizing the (conditional) expected average reward while guaranteeing an optimal probability of staying safe (w.r.t. some safe set of states). For the first combination of objectives, which covers the classical frozen lake environment from reinforcement learning, we also report on experiments performed using a prototype implementation of our algorithm and compare it with what can be obtained from state-of-the-art probabilistic model checkers solving optimal reachability.

cs.GT↗

Exact generalized Turán number for $K_3$ versus suspension of $P_4$

Let $P_4$ denote the path graph on $4$ vertices. The suspension of $P_4$, denoted by $\widehat P_4$, is the graph obtained via adding an extra vertex and joining it to all four vertices of $P_4$. In this note, we demonstrate that for $n\ge 8$, the maximum number of triangles in any $n$-vertex graph not containing $\widehat P_4$ is $\left\lfloor n^2/8\right\rfloor$. Our method uses simple induction along with computer programming to prove a base case of the induction hypothesis.

math.CO↗

Representing Fields without Correspondences: the Lifted Euler Characteristic Transform

Topological transforms have been very useful in statistical analysis of shapes or surfaces without restrictions that the shapes are diffeomorphic and requiring the estimation of correspondence maps. In this paper we introduce two topological transforms that generalize from shapes to fields, $f:\mathbf{R}^3 \rightarrow \mathbf{R}$. Both transforms take a field and associate to each direction $v\in S^{d-1}$ a summary obtained by scanning the field in the direction $v$. The transforms we introduce are of interest for both applications as well as their theoretical properties. The topological transforms for shapes are based on an Euler calculus on sets. A key insight in this paper is that via a lifting argument one can develop an Euler calculus on real valued functions from the standard Euler calculus on sets, this idea is at the heart of the two transforms we introduce. We prove the transforms are injective maps. We show for particular moduli spaces of functions we can upper bound the number of directions needed determine any particular function.

math.AT↗

A Sheaf-Theoretic Construction of Shape Space

We present a sheaf-theoretic construction of shape space -- the space of all shapes. We do this by describing a homotopy sheaf on the poset category of constructible sets, where each set is mapped to its Persistent Homology Transform (PHT). Recent results that build on fundamental work of Schapira have shown that this transform is injective, thus making the PHT a good summary object for each shape. Our homotopy sheaf result allows us to "glue" PHTs of different shapes together to build up the PHT of a larger shape. In the case where our shape is a polyhedron we prove a generalized nerve lemma for the PHT. Finally, by re-examining the sampling result of Smale-Niyogi-Weinberger, we show that we can reliably approximate the PHT of a manifold by a polyhedron up to arbitrary precision.

math.AT↗

Large Deviation Asymptotics and Bayesian Posterior Consistency on Stochastic Processes and Dynamical Systems

We consider generalized Bayesian inference on stochastic processes and dynamical systems with potentially long-range dependency. Given a sequence of observations, a class of parametrized model processes with a prior distribution, and a loss function, we specify the generalized posterior distribution. The problem of frequentist posterior consistency is concerned with whether as more and more samples are observed, the posterior distribution on parameters will asymptotically concentrate on the "right" parameters. We show that posterior consistency can be derived using a combination of classical large deviation techniques, such as Varadhan's lemma, conditional/quenched large deviations, annealed large deviations, and exponential approximations. We show that the posterior distribution will asymptotically concentrate on parameters that minimize the expected loss and a divergence term, and we identify the divergence term as the Donsker-Varadhan relative entropy rate from process-level large deviations. As an application, we prove new quenched and annealed large deviation asymptotics and new Bayesian posterior consistency results for a class of mixing stochastic processes. In the case of Markov processes, one can obtain explicit conditions for posterior consistency, whenever estimates for log-Sobolev constants are available, which makes our framework essentially a black box. We also recover state-of-the-art posterior consistency on classical dynamical systems with a simple proof. Our approach has the potential of proving posterior consistency for a wide range of Bayesian procedures in a unified way.

math.ST↗

Concentration inequalities and optimal number of layers for stochastic deep neural networks

We state concentration inequalities for the output of the hidden layers of a stochastic deep neural network (SDNN), as well as for the output of the whole SDNN. These results allow us to introduce an expected classifier (EC), and to give probabilistic upper bound for the classification error of the EC. We also state the optimal number of layers for the SDNN via an optimal stopping procedure. We apply our analysis to a stochastic version of a feedforward neural network with ReLU activation function.

cs.LG↗

Triangles in graphs without bipartite suspensions

Given graphs $T$ and $H$, the generalized Turán number ex$(n,T,H)$ is the maximum number of copies of $T$ in an $n$-vertex graph with no copies of $H$. Alon and Shikhelman, using a result of Erd\H os, determined the asymptotics of ex$(n,K_3,H)$ when the chromatic number of $H$ is greater than 3 and proved several results when $H$ is bipartite. We consider this problem when $H$ has chromatic number 3. Even this special case for the following relatively simple 3-chromatic graphs appears to be challenging. The suspension $\widehat H$ of a graph $H$ is the graph obtained from $H$ by adding a new vertex adjacent to all vertices of $H$. We give new upper and lower bounds on ex$(n,K_3,\widehat{H})$ when $H$ is a path, even cycle, or complete bipartite graph. One of the main tools we use is the triangle removal lemma, but it is unclear if much stronger statements can be proved without using the removal lemma.

math.CO↗

Global Optimality of Elman-type RNN in the Mean-Field Regime

We analyze Elman-type Recurrent Reural Networks (RNNs) and their training in the mean-field regime. Specifically, we show convergence of gradient descent training dynamics of the RNN to the corresponding mean-field formulation in the large width limit. We also show that the fixed points of the limiting infinite-width dynamics are globally optimal, under some assumptions on the initialization of the weights. Our results establish optimality for feature-learning with wide RNNs in the mean-field regime

stat.ML↗

Regularized Bayesian best response learning in finite games

We introduce the notion of regularized Bayesian best response (RBBR) learning dynamic in heterogeneous population games. We obtain such a dynamic via perturbation by an arbitrary lower semicontinuous, strongly convex regularizer in Bayesian population games with finitely many strategies. We provide a sufficient condition for the existence of rest points of the RBBR learning dynamic, and hence the existence of regularized Bayesian equilibrium in Bayesian population games. These equilibria are shown to approximate the original Bayesian equilibria for vanishingly small perturbations. We also explore the fundamental properties of the RBBR learning dynamic, which includes the existence of unique continuous solutions from arbitrary initial conditions, as well as the continuity of the solution trajectories thus obtained with respect to the initial conditions. Finally, as application the the theory we introduce the notions of Bayesian potential and Bayesian negative semidefinite games and provide convergence results for such games.

math.OC↗

Extended probabilities and their application to statistical inference

We propose a new, more general definition of extended probability measures. We study their properties and provide a behavioral interpretation. We put them to use in an inference procedure, whose environment is canonically represented by the probability space $(Ω,\mathcal{F},P)$, when both $P$ and the composition of $Ω$ are unknown. We develop an ex ante analysis -- taking place before the statistical analysis requiring knowledge of $Ω$ -- in which the true composition of $Ω$ is progressively learned. We describe how to update extended probabilities in this setting, and introduce the concept of lower extended probabilities. We apply our findings to a species sampling problem and to the study of the boomerang effect (the empirical observation that sometimes persuasion yields the opposite effect: the persuaded agent moves their opinion away from the opinion of the persuading agent).

math.ST↗

Ergodic Theorems for Dynamic Imprecise Probability Kinematics

We formulate an ergodic theory for the (almost sure) limit $\mathcal{P}^\text{co}_{\tilde{\mathcal{E}}}$ of a sequence $(\mathcal{P}^\text{co}_{\mathcal{E}_n})$ of successive dynamic imprecise probability kinematics (DIPK, introduced in Caprio and Gong, 2021) updates of a set $\mathcal{P}^\text{co}_{\mathcal{E}_0}$ representing the initial beliefs of an agent. As a consequence, we formulate a strong law of large numbers.

math.ST↗

Multiple testing with persistent homology

In this paper we propose a computationally efficient multiple hypothesis testing procedure for persistent homology. The computational efficiency of our procedure is based on the observation that one can empirically simulate a null distribution that is universal across many hypothesis testing applications involving persistence homology. Our observation suggests that one can simulate the null distribution efficiently based on a small number of summaries of the collected data and use this null in the same way that p-value tables were used in classical statistics. To illustrate the efficiency and utility of the null distribution we provide procedures for rejecting acyclicity with both control of the Family-Wise Error Rate (FWER) and the False Discovery Rate (FDR). We will argue that the empirical null we propose is very general conditional on a few summaries of the data based on simulations and limit theorems for persistent homology for point processes.

cs.CG↗

Tight query complexity bounds for learning graph partitions

Given a partition of a graph into connected components, the membership oracle asserts whether any two vertices of the graph lie in the same component or not. We prove that for $n\ge k\ge 2$, learning the components of an $n$-vertex hidden graph with $k$ components requires at least $(k-1)n-\binom k2$ membership queries. Our result improves on the best known information-theoretic bound of $Ω(n\log k)$ queries, and exactly matches the query complexity of the algorithm introduced by [Reyzin and Srivastava, 2007] for this problem. Additionally, we introduce an oracle, with access to which one can learn the number of components of $G$ in asymptotically fewer queries than learning the full partition, thus answering another question posed by the same authors. Lastly, we introduce a more applicable version of this oracle, and prove asymptotically tight bounds of $\widetildeΘ(m)$ queries for both learning and verifying an $m$-edge hidden graph $G$ using it.

cs.LG↗

A Grover search-based algorithm for the list coloring problem

Graph coloring is a computationally difficult problem, and currently the best known classical algorithm for $k$-coloring of graphs on $n$ vertices has runtimes $Ω(2^n)$ for $k\ge 5$. The list coloring problem asks the following more general question: given a list of available colors for each vertex in a graph, does it admit a proper coloring? We propose a quantum algorithm based on Grover search to quadratically speed up exhaustive search. Our algorithm loses in complexity to classical ones in specific restricted cases, but improves exhaustive search for cases where the lists and graphs considered are arbitrary in nature.

quant-ph↗

A Methodology for Exploring Deep Convolutional Features in Relation to Hand-Crafted Features with an Application to Music Audio Modeling

Understanding the features learned by deep models is important from a model trust perspective, especially as deep systems are deployed in the real world. Most recent approaches for deep feature understanding or model explanation focus on highlighting input data features that are relevant for classification decisions. In this work, we instead take the perspective of relating deep features to well-studied, hand-crafted features that are meaningful for the application of interest. We propose a methodology and set of systematic experiments for exploring deep features in this setting, where input feature importance approaches for deep feature understanding do not apply. Our experiments focus on understanding which hand-crafted and deep features are useful for the classification task of interest, how robust these features are for related tasks and how similar the deep features are to the meaningful hand-crafted features. Our proposed method is general to many application areas and we demonstrate its utility on orchestral music audio data.

cs.SD↗