SearcharxivSearch

arXiv subjects

Bhaswar B. Bhattacharya

Publications and source records attributed to Bhaswar B. Bhattacharya.

At least 19 recordsLinked to original sources

Large Planar Point Sets Contain 4 Collinear Points or Almost 7-Cliques, and Related Results

We prove that every sufficiently large finite planar point set contains either four collinear points or seven points with at most one non-visible pair. More generally, we show that for every fixed graph $H$ with chromatic number at most five, or with chromatic number six and a color-critical edge, the visibility graph of every sufficiently large finite planar point set with no four collinear points contains a copy of $H$. These results extend the recent breakthrough of Bonnet (2026), guaranteeing six pairwise visible points, and come within one visibility edge of the next open case of the big-line-big-clique conjecture.

math.CO

Almost Empty Monochromatic Triangles With Many Colors

Given integers $c\geq 2$ and $s\geq 0$, let $\mathsf{M}_3(c,s)$ denote the least integer such that every set of at least $\mathsf{M}_3(c,s)$ points in the plane, no three on a line, colored with $c$ colors, contains a monochromatic triangle with at most $s$ interior points. Further, let $λ_3(c)$ be the least integer such that $\mathsf{M}_3(c,λ_3(c))<\infty$. \citet{colorempty} proved that, for every $c\geq 2$, $$\left\lfloor\frac{c-1}{2}\right\rfloor \leq λ_3(c)\leq c-2.$$ Later, \citet{cravioto2019almost} improved the upper bound to $c-3$, for $c\geq 4$. In this paper, we refine their argument to obtain the following asymptotic improvement: $$λ_3(c) \leq c-\sqrt{c\log c}+o (\sqrt{c\log c} ),$$ for all sufficiently large $c$. We also show that every $c$-coloring of a sufficiently large Horton set contains a monochromatic triangle with at most $\lfloor \frac{c-1}{2} \rfloor$ interior points. This shows that the aforementioned lower bound on $λ_3(c)$ is sharp within the class of Horton sets. We conclude with a conjecture on the large-color asymptotics of $λ_3(c)$.

math.CO

Conditional Mean Independence, Global Sensitivity Analysis, and Variable Screening Using nearest neighbor Graphs

Quantifying how much of the variation in a response is explained by its conditional mean is central to many statistical problems. In this paper, we develop a unified framework for this problem, centered on a normalized conditional mean discrepancy that coincides with the classical Sobol' index for univariate responses and its natural trace-based extension for multivariate responses. We propose a simple estimator of this measure based on nearest-neighbor graphs that can be computed in near-linear time. We establish its consistency and derive its rate of convergence. Further, under the null hypothesis of conditional mean independence, a studentized version of the statistic is asymptotically standard normal. This leads to a computationally efficient test for conditional mean independence that attains the correct asymptotic level and is universally consistent, without requiring bootstrap calibration or sample splitting. Building on the same estimator, we develop a model-free sequential variable-screening procedure that selects a sufficient set with high probability and an exponential error bound. We also discuss extensions of the framework to quantifying interaction effects through higher-order Sobol' indices. Simulations and real-data experiments demonstrate strong power, effective variable screening, and substantial computational gains over existing methods.

stat.ME

Thresholds and Fluctuations for Colorful Arithmetic Progressions in Sparse Random Colorings

In this paper, we derive thresholds and fluctuations for arithmetic progressions with prescribed color patterns in sparse random colorings of $[n]:=\{1, 2, \ldots, n\}$, where each element of $[n]$ is colored independently according to a given probability vector. For any admissible ordered palette of colors, we determine the full multi-parameter threshold region for the appearance of a colorful arithmetic progression. The threshold is governed by two competing mechanisms: a global first-moment condition and a local color-availability condition, resulting in a polyhedral satisfiability region, with a piecewise-polyhedral threshold surface. In the satisfiability region we establish asymptotic normality for the number of colorful arithmetic progressions of a given length, with an explicit rate of convergence in Wasserstein distance. On the threshold surface, we identify three distinct asymptotic regimes: Poisson, compound Poisson with mixed Poisson jumps, and compound Poisson with uniform jumps, after an appropriate normalization. These results provide a complete description of the threshold and fluctuation behavior of general colored arithmetic progressions under sparse random colorings, in a unified framework that interpolates between classical uncolored/monochromatic progressions in binomial random subsets and multicolored, including rainbow, arithmetic progressions.

math.CO

On the Number of Almost Empty Monochromatic Triangles

In this paper, we consider the problem of counting almost empty monochromatic triangles in colored planar point sets, that is, triangles whose vertices are all assigned the same color and that contain only a few interior points. Specifically, we show that any $c$-coloring of a set of $n$ points in the plane in general position (that is, no three on a line) contains $Ω(n^2)$ monochromatic triangles with at most $c-1$ interior points and $Ω(n^{\frac{4}{3}})$ monochromatic triangles with at most $c-2$ interior points, for any fixed $c \geq 2$. The latter, in particular, generalizes the result of Pach and Tóth (2013) on the number of monochromatic empty triangles in 2-colored point sets, to the setting of multiple colors and monochromatic triangles with a few interior points. We also derive the limiting value of the expected number of triangles with $s$ interior points in random point sets, for any integer $s \geq 0$. As a result, we obtain the expected number of monochromatic triangles with at most $s$ interior points in random colorings of random point sets.

math.CO

Colorful Exponential Random Graph Models

In this paper, we initiate the study of colored exponential random graph models (ERGMs), a class of exponential-family models for networks with multiple types of edge relations. Using the framework of probability graphons, we first derive a variational representation for the limiting free energy, whose maximizers determine the asymptotic structure of typical samples from the model. Then we identify several general families of colored ERGMs exhibiting replica symmetry, where the variational problem has constant maximizers and the model asymptotically concentrates on product colorings with independent edges. For general colored ERGMs, we derive Euler-Lagrange fixed-point equations for the variational maximizers, which in turn yield a general high-temperature uniqueness criterion. In the complementary zero-temperature regime, we establish a two-level selection principle: the leading energy term determines the ground states, while the lower-order energy terms, combined with entropy, act as a tie-breaker to determine the asymptotic zero-temperature structure of the model. We illustrate this principle through the induced wedge and rainbow triangle ERGMs. Both models have natural interpretations in multitype networks, and their zero-temperature limits exhibit interesting structures that connect to well-known results in extremal combinatorics. We further establish finite-temperature symmetry breaking for both these models and complement the rigorous results with numerical experiments.

math.PR

Thresholds and Fluctuations of Submultiplexes in Random Multiplex Networks

In a multiplex network a common set of nodes is connected through different types of interactions, each represented as a separate graph (layer) within the network. In this paper, we study the asymptotic properties of submultiplexes, the counterparts of subgraphs (motifs) in single-layer networks, in the correlated Erdős-Rényi multiplex model. This is a random multiplex model with two layers, where the graphs in each layer marginally follow the classical (single-layer) Erdős-Rényi model, while the edges across layers are correlated. We derive the precise threshold condition for the emergence of a fixed submultiplex $\boldsymbol{H}$ in a random multiplex sampled from the correlated Erdős-Rényi model. Specifically, we show that the satisfiability region, the regime where the random multiplex contains infinitely many copies of $\boldsymbol{H}$, forms a polyhedral subset of $\mathbb{R}^3$. Furthermore, within this region the count of $\boldsymbol{H}$ is asymptotically normal, with an explicit convergence rate in the Wasserstein distance. We also establish various Poisson approximation results for the count of $\boldsymbol{H}$ on the boundary of the threshold, which depends on a notion of balance of submultiplexes. Collectively, these results provide an asymptotic theory for small submultiplexes in the correlated multiplex model, analogous to the classical theory of small subgraphs in random graphs.

math.PR

Ising Models on Inhomogeneous Random Graphs: Inference, Local Asymptotic Minimaxity, and Limit of Experiments

In this paper, we develop an inferential framework with sharp asymptotic optimality guarantees for Ising models on inhomogeneous random graphs in the subcritical parameter regime. We begin by characterizing the asymptotic distribution of the maximum likelihood (ML) estimate of the natural parameter, based on a single sample from the underlying model, covering both sparse and dense network regimes. Next, to overcome the computational intractability of the ML method, we propose a simple closed-form estimate obtained from a one-step approximation to the likelihood equation. We show that this estimate attains the same asymptotic distribution and variance as the ML estimate, thereby yielding a computationally efficient and asymptotically valid confidence interval for the natural parameter. We complement these inferential results by establishing a Hájek--Le Cam-type local asymptotic minimax theorem, showing that the proposed estimate achieves the smallest possible asymptotic maximum risk, both in rate and in leading constant, over shrinking neighborhoods of the true parameter. We also derive the corresponding limit of experiments. To the best of our knowledge, these are among the first sharp asymptotic optimality results for network-dependent data. Finally, we study goodness-of-fit testing for the natural parameter, deriving the local power of the likelihood ratio test and minimax detection rates. Our analysis relies on new fluctuation results for the sufficient statistic (Hamiltonian) and for the random partition function of Ising models on inhomogeneous random graphs, which are of independent interest.

math.ST

Transitivity in Inhomogeneous Random Tournaments

Paired-comparison data are naturally represented by tournaments, where transitivity corresponds to the existence of a global ranking consistent with all pairwise outcomes. Accordingly, the classical Kendall-Smith coefficient of consistency measures deviations from transitivity in a tournament by counting the number of circular triads (directed $3$-cycles). In this paper, we characterize the fluctuations of the number of circular triads in inhomogeneous random tournaments and develop an inferential framework for the consistency coefficient. Specifically, we consider the $W$-random tournament model, where the comparison probabilities are determined by a tournamenton $W$, the analogue of a graphon in the tournament setting. We show that, for a $W$-random tournament on $n$ vertices, the number of circular triads exhibits three different fluctuation regimes, determined by suitable notions of regularity and uniformity of $W$. We further develop a novel tournamenton multiplier bootstrap that consistently approximates the limiting distribution of the circular-triad count in the relevant asymptotic regime. Combining this with procedures for testing regularity and uniformity, we design an algorithm for constructing confidence intervals for the consistency coefficient that is asymptotically valid for all tournamentons. We also obtain structural characterizations of tournamentons for which the limiting distribution of the number of circular triads exhibits specific degeneracies. These results can also be viewed through the lens of tournament quasirandomness and may be of independent interest.

math.PR

Asymptotic Normality of Subgraph Counts in Sparse Inhomogeneous Random Graphs

In this paper, we derive the asymptotic distribution of the number of copies of a fixed graph $H$ in a random graph $G_n$ sampled from a sparse graphon model. Specifically, we provide a refined analysis that separates the contributions of edge randomness and vertex-label randomness, allowing us to identify distinct sparsity regimes in which each component dominates or both contribute jointly to the fluctuations. As a result, we establish asymptotic normality for the count of any fixed graph $H$ in $G_n$ across the entire range of sparsity (above the containment threshold for $H$ in $G_n$). These results provide a complete description of subgraph count fluctuations in sparse inhomogeneous networks, closing several gaps in the existing literature that were limited to specific motifs or suboptimal sparsity assumptions.

math.PR

Multiplexons: Limits of Multiplex Networks

In a multiplex network, a set of nodes is connected by different types of interactions, each represented as a separate layer within the network. Multiplexes have emerged as a key instrument for modeling large-scale complex systems, due to the widespread coexistence of diverse interactions in social, industrial, and biological domains. This motivates the development of a rigorous and readily applicable framework for studying properties of large multiplex networks. In this article, we provide a self-contained introduction to the limit theory of dense multiplex networks, analogous to the theory of graphons (limit theory of dense graphs). As applications, we derive limiting analogues of commonly used multiplex features, such as degree distributions and clustering coefficients. We also present a range of illustrative examples, including correlated versions of Erdős-Rényi and inhomogeneous random graph models and dynamic networks. Finally, we discuss how multiplex networks fit within the broader framework of decorated graphs, and how the convergence results can be recovered from the limit theory of decorated graphs. Several future directions are outlined for further developing the multiplex limit theory.

math.PR

Monochromatic Subgraphs in Randomly Colored Dense Multiplex Networks

Given a sequence of graphs $G_n$ and a fixed graph $H$, denote by $T(H, G_n)$ the number of monochromatic copies of the graph $H$ in a uniformly random $c$-coloring of the vertices of $G_n$. In this paper we study the joint distribution of a finite collection of monochromatic graph counts in networks with multiple layers (multiplex networks). Specifically, given a finite collection of graphs $H_1, H_2, \ldots, H_d$ we derive the joint distribution of $(T(H_1, G_n^{(1)}), T(H_2, G_n^{(2)}), \ldots, T(H_d, G_n^{(d)}))$, where $\boldsymbol{G}_n = (G_n^{(1)}, G_n^{(2)}, \ldots, G_n^{(d)})$ is a collection of dense graphs on the same vertex set converging in the joint cut-metric. The limiting distribution is the sum of 2 independent components: a multivariate Gaussian and a sum of independent bivariate stochastic integrals. This extends previous results on the marginal convergence of monochromatic subgraphs in a sequence of graphs to the joint convergence of a finite collection of monochromatic subgraphs in a sequence of multiplex networks. Several applications and examples are discussed.

math.PR

Motif Estimation via Subgraph Sampling: The Fourth Moment Phenomenon

Network sampling is an indispensable tool for understanding features of large complex networks where it is practically impossible to search over the entire graph. In this paper, we develop a framework for statistical inference for counting network motifs, such as edges, triangles, and wedges, in the widely used subgraph sampling model, where each vertex is sampled independently, and the subgraph induced by the sampled vertices is observed. We derive necessary and sufficient conditions for the consistency and the asymptotic normality of the natural Horvitz-Thompson (HT) estimator, which can be used for constructing confidence intervals and hypothesis testing for the motif counts based on the sampled graph. In particular, we show that the asymptotic normality of the HT estimator exhibits an interesting fourth-moment phenomenon, which asserts that the HT estimator (appropriately centered and rescaled) converges in distribution to the standard normal whenever its fourth-moment converges to 3 (the fourth-moment of the standard normal distribution). As a consequence, we derive the exact thresholds for consistency and asymptotic normality of the HT estimator in various natural graph ensembles, such as sparse graphs with bounded degree, Erdos-Renyi random graphs, random regular graphs, and dense graphons.

math.ST

Joint Poisson Convergence of Monochromatic Hyperedges in Multiplex Hypergraphs

Given a sequence of $r$-uniform hypergraphs $H_n$, denote by $T(H_n)$ the number of monochromatic hyperedges when the vertices of $H_n$ are colored uniformly at random with $c = c_n$ colors. In this paper, we study the joint distribution of monochromatic hyperedges for hypergraphs with multiple layers (multiplex hypergraphs). Specifically, we consider the joint distribution of ${\bf T} _n:= (T(H_n^{(1)}), T(H_n^{(2)}))$, for two sequences of hypergraphs $H_n^{(1)}$ and $H_n^{(2)}$ on the same set of vertices. We will show that the joint distribution of ${\bf T}_n$ converges to (possibly dependent) Poisson distributions whenever the mean vector and the covariance matrix of ${\bf T}_n$ converge. In other words, the joint Poisson approximation of ${\bf T}_n$ is determined only by the convergence of its first two moments. This generalizes recent results on the second moment phenomenon for Poisson approximation from graph coloring to hypergraph coloring and from marginal convergence to joint convergence. Applications include generalizations of the birthday problem, counting monochromatic subgraphs in randomly colored graphs, and counting monochromatic arithmetic progressions in randomly colored integers. Extensions to random hypergraphs and weighted hypergraphs are also discussed.

math.PR

High Dimensional Logistic Regression Under Network Dependence

Logistic regression is key method for modeling the probability of a binary outcome based on a collection of covariates. However, the classical formulation of logistic regression relies on the independent sampling assumption, which is often violated when the outcomes interact through an underlying network structure, such as over a temporal/spatial domain or on a social network. This necessitates the development of models that can simultaneously handle both the network `peer-effect' and the effect of high-dimensional covariates. In this paper, we develop a framework for incorporating such dependencies in a high-dimensional logistic regression model by introducing a quadratic interaction term, as in the Ising model, designed to capture the pairwise interactions from the underlying network. The resulting model can also be viewed as an Ising model, where the node-dependent external fields linearly encode the high-dimensional covariates. We propose a penalized maximum pseudo-likelihood method for estimating the network peer-effect and the effect of the covariates (the regression coefficients), which, in addition to handling the high-dimensionality of the parameters, conveniently avoids the computational intractability of the maximum likelihood approach. Under various standard regularity conditions, we show that the corresponding estimate attains the classical high-dimensional rate of consistency. Our results imply that even under network dependence it is possible to consistently estimate the model parameters at the same rate as in classical (independent) logistic regression, when the true parameter is sparse and the underlying network is not too dense. We also develop an efficient algorithm for computing the estimates and validate our theoretical results in numerical experiments. An application to selecting genes in clustering spatial transcriptomics data is also discussed.

math.ST

Higher-Order Graphon Theory: Fluctuations, Degeneracies, and Inference

Exchangeable random graphs, which include some of the most widely studied network models, have emerged as the mainstay of statistical network analysis in recent years. Graphons, which are the central objects in graph limit theory, provide a natural way to sample exchangeable random graphs. It is well known that network moments (motif/subgraph counts) identify a graphon (up to an isomorphism), hence, understanding the sampling distribution of subgraph counts in random graphs sampled from a graphon is pivotal for nonparametric network inference. In this paper, we derive the joint asymptotic distribution of any finite collection of network moments in random graphs sampled from a graphon, that includes both the non-degenerate case (where the distribution is Gaussian) as well as the degenerate case (where the distribution has both Gaussian or non-Gaussian components). This provides the higher-order fluctuation theory for subgraph counts in the graphon model. We also develop a novel multiplier bootstrap for graphons that consistently approximates the limiting distribution of the network moments (both in the Gaussian and non-Gaussian regimes). Using this and a procedure for testing degeneracy, we construct joint confidence sets for any finite collection of motif densities. This provides a general framework for statistical inference based on network moments in the graphon model. To illustrate the broad scope of our results we also consider the problem of detecting global structure (that is, testing whether the graphon is a constant function) based on small subgraphs. We propose a consistent test for this problem, invoking celebrated results on quasi-random graphs, and derive its limiting distribution both under the null and the alternative.

math.ST

Fluctuations of Quadratic Chaos

In this paper we characterize all distributional limits of the random quadratic form $T_n =\sum_{1\le u< v\le n} a_{u, v} X_u X_v$, where $((a_{u, v}))_{1\le u,v\le n}$ is a $\{0, 1\}$-valued symmetric matrix with zeros on the diagonal and $X_1, X_2, \ldots, X_n$ are i.i.d.~ mean $0$ variance $1$ random variables with common distribution function $F$. In particular, we show that any distributional limit of $S_n:=T_n/\sqrt{\mathrm{Var}[T_n]}$ can be expressed as the sum of three independent components: a Gaussian, a (possibly) infinite weighted sum of independent centered chi-squares, and a Gaussian mixture with a random variance. As a consequence, we prove a fourth moment theorem for the asymptotic normality of $S_n$, which applies even when $F$ does not have finite fourth moment. More formally, we show that $S_n$ converges to $N(0, 1)$ if and only if the fourth moment of $S_n$ (appropriately truncated when $F$ does not have finite fourth moment) converges to 3 (the fourth moment of the standard normal distribution).

math.PR

Boosting the Power of Kernel Two-Sample Tests

The kernel two-sample test based on the maximum mean discrepancy (MMD) is one of the most popular methods for detecting differences between two distributions over general metric spaces. In this paper we propose a method to boost the power of the kernel test by combining MMD estimates over multiple kernels using their Mahalanobis distance. We derive the asymptotic null distribution of the proposed test statistic and use a multiplier bootstrap approach to efficiently compute the rejection region. The resulting test is universally consistent and, since it is obtained by aggregating over a collection of kernels/bandwidths, is more powerful in detecting a wide range of alternatives in finite samples. We also derive the distribution of the test statistic for both fixed and local contiguous alternatives. The latter, in particular, implies that the proposed test is statistically efficient, that is, it has non-trivial asymptotic (Pitman) efficiency. The consistency properties of the Mahalanobis and other natural aggregation methods are also explored when the number of kernels is allowed to grow with the sample size. Extensive numerical experiments are performed on both synthetic and real-world datasets to illustrate the efficacy of the proposed method over single kernel tests. The computational complexity of the proposed method is also studied, both theoretically and in simulations. Our asymptotic results rely on deriving the joint distribution of MMD estimates using the framework of multiple stochastic integrals, which is more broadly useful, specifically, in understanding the efficiency properties of recently proposed adaptive MMD tests based on kernel aggregation and also in developing more computationally efficient (linear time) tests that combine multiple kernels. We conclude with an application of the Mahalanobis aggregation method for kernels with diverging scaling parameters.

stat.ME