SearcharxivSearch

arXiv subjects

Mingao Yuan

Publications and source records attributed to Mingao Yuan.

At least 19 recordsLinked to original sources

On the number of small edge-weighted subgraphs

Subgraph counting is a fundamental task that underpins several network analysis methodologies, including community detection and graph two-sample tests. Counting subgraphs is a computationally intensive problem. Substantial research has focused on developing efficient algorithms and strategies to make it feasible for larger unweighted graphs. Implementing those algorithms can be a significant hurdle for data professionals or researchers with limited expertise in algorithmic principles and programming. Furthermore, many real-world networks are weighted. Computing the number of weighted subgraphs in weighted networks presents a computational challenge, as no efficient algorithm exists for the worst-case scenario. In this paper, we derive explicit formulas for counting small edge-weighted subgraphs using the weighted adjacency matrix. These formulas are applicable to unweighted networks, offering a simple and highly practical analytical tool for researchers across various scientific domains. In addition, we introduce a generalized methodology for calculating arbitrary weighted subgraphs.

math.CO

The interplay between network transitivity and community structure

Recent empirical observations suggest that network transitivity is highly correlated with community structure in many real-world networks. In this paper, we theoretically investigate this relationship by deriving the limits of the global and average clustering coefficients for the geometric block model (GBM). Both limits exhibit a phase transition; specifically, the functional forms of the limit functions differ between the weak and strong community structure strength regimes. For a GBM with balanced communities, the limits of the global and average clustering coefficients are identical, whereas these limits differ for unbalanced communities. In general, the clustering coefficients do not exhibit a monotonic relationship with community structure strength. Particularly, for a balanced GBM where the within-community edge probability is a constant multiple of the between-community edge probability, the limit decreases from $3/4$ to $3/5$ and subsequently increases toward an asymptotic upper bound of $3/4$ as the multiple grows from one. A similar pattern is observed for the global clustering coefficient in unbalanced settings, where both limits exhibit an explicit dependence on community size.

math.ST

Central limit theorem for the global clustering coefficient of random geometric graphs

The global clustering coefficient serves as a powerful metric for the structural analysis and comparison of complex networks. Random geometric graphs offer a realistic framework for representing the spatial constraints and geometry often found in real-world network datasets. In this paper, we establish a central limit theorem for the global clustering coefficient of random geometric graphs. Our main result identifies the centering and scaling sequences required for convergence in law to the standard normal distribution. Our approach varies by regime: in the dense case, we employ the Lyapunov CLT; in the intermediate case, we utilize the asymptotic theory of $U$-statistics with sample-size-dependent kernels; and in the sparse regime, we use the method of moments to derive the asymptotic distribution. Notably, the convergence rates for non-uniform and uniform random geometric graphs diverge in the dense regime, yet they coincide in the sparse regime. In addition, we find that the global clustering coefficient for both uniform and non-uniform RGGs is asymptotically equal to $3/4$

math.ST

The weak law of large numbers for the friendship paradox index

The friendship paradox index is a network summary statistic used to quantify the friendship paradox, which describes the tendency for an individual's friends to have more friends than the individual. In this paper, we utilize Markov's inequality to derive the weak law of large numbers for the friendship paradox index in a random geometric graph, a widely-used model for networks with spatial dependence and geometry. For uniform random geometric graph, where the nodes are uniformly distributed in a space, the friendship paradox index is asymptotically equal to $1/4$. On the contrary, in nonuniform random geometric graphs, the nonuniform node distribution leads to distinct limiting properties for the index. In the relatively sparse regime, the friendship paradox index is still asymptotically equal to $1/4$, the same as in the uniform case. In the intermediate sparse regime, however, the index converges in probability to $1/4$ plus a constant that is explicitly dependent on the node distribution. Finally, in the relatively dense case, the index diverges to infinity as the graph size increases. Our results highlight the sharp contrast between the uniform case and its nonuniform counterpart.

math.ST

Hypothesis testing for the dimension of random geometric graph

Random geometric graphs (RGGs) offer a powerful tool for analyzing the geometric and dependence structures in real-world networks. For example, it has been observed that RGGs are a good model for protein-protein interaction networks. In RGGs, nodes are randomly distributed over an $m$-dimensional metric space, and edges connect the nodes if and only if their distance is less than some threshold. When fitting RGGs to real-world networks, the first step is probably to input or estimate the dimension $m$. However, it is not clear whether the prespecified dimension is equal to the true dimension. In this paper, we investigate this problem using hypothesis testing. Under the null hypothesis, the dimension is equal to a specific value, while the alternative hypothesis asserts the dimension is not equal to that value. We propose the first statistical test. Under the null hypothesis, the proposed test statistic converges in law to the standard normal distribution, and under the alternative hypothesis, the test statistic is unbounded in probability. We derive the asymptotic distribution by leveraging the asymptotic theory of degenerate U-statistics with kernel function dependent on the number of nodes. This approach differs significantly from prevailing methods used in network hypothesis testing problems. Moreover, we also propose an efficient approach to compute the test statistic based on the adjacency matrix. Simulation studies show that the proposed test performs well. We also apply the proposed test to multiple real-world networks to test their dimensions.

stat.ME

Hypothesis testing for the uniformity of random geometric graph

Random geometric graphs are widely used in modeling geometry and dependence structure in networks. In a random geometric graph, nodes are independently generated from some probability distribution $F$ over a metric space, and edges link nodes if their distance is less than some threshold. Most studies assume the distribution $F$ to be uniform. However, recent research shows that some real-world networks may be better modeled by nonuniform distribution $F$. Moreover, graphs with nonuniform $F$ have notably different properties from graphs with uniform $F$. A fundamental question is: given a network from a random geometric graph, is the distribution $F$ uniform or not? In this paper, we approach this question through hypothesis testing. This problem is particularly challenging due to the inherent dependencies among edges in random geometric graphs, a property not present in classic random graphs. We propose the first statistical test. Under the null hypothesis, the test statistic converges in distribution to the standard normal distribution. The asymptotic distribution is derived using the asymptotic theory of degenerate U-statistics with a kernel function dependent on the number of nodes. This technique is different from existing methods in network hypothesis testing problems. In addition, we present a method for efficiently calculating the test statistic directly from the adjacency matrix. We also analytically characterize the power of the proposed test. The simulation study shows that the proposed uniformity test has high power. Real data applications are also provided.

stat.ME

Asymptotic distribution of the global clustering coefficient in a random annulus graph

The global clustering coefficient is an effective measure for analyzing and comparing the structures of complex networks. The random annulus graph is a modified version of the well-known Erdős-Rényi random graph. It has been recently proposed in modeling network communities. This paper investigates the asymptotic distribution of the global clustering coefficient in a random annulus graph. It is demonstrated that the standardized global clustering coefficient converges in law to the standard normal distribution. The result is established using the asymptotic theory of degenerate U-statistics with a sample-size dependent kernel. As far as we know, this method is different from established approaches for deriving asymptotic distributions of network statistics. Moreover, we get the explicit expression of the limit of the global clustering coefficient.

stat.ME

Testing common invariant subspace of multilayer networks

Graph (or network) is a mathematical structure that has been widely used to model relational data. As real-world systems get more complex, multilayer (or multiple) networks are employed to represent diverse patterns of relationships among the objects in the systems. One active research problem in multilayer networks analysis is to study the common invariant subspace of the networks, because such common invariant subspace could capture the fundamental structural patterns and interactions across all layers. Many methods have been proposed to estimate the common invariant subspace. However, whether real-world multilayer networks share the same common subspace remains unknown. In this paper, we first attempt to answer this question by means of hypothesis testing. The null hypothesis states that the multilayer networks share the same subspace, and under the alternative hypothesis, there exist at least two networks that do not have the same subspace. We propose a Weighted Degree Difference Test, derive the limiting distribution of the test statistic and provide an analytical analysis of the power. Simulation study shows that the proposed test has satisfactory performance, and a real data application is provided.

stat.ME

Central limit theorem for the average closure coefficient

Many real-world networks exhibit the phenomenon of edge clustering, which is typically measured by the average clustering coefficient. Recently, an alternative measure, the average closure coefficient, is proposed to quantify local clustering. It is shown that the average closure coefficient possesses a number of useful properties and can capture complementary information missed by the classical average clustering coefficient. In this paper, we study the asymptotic distribution of the average closure coefficient of a heterogeneous Erdös-Rényi random graph. We prove that the standardized average closure coefficient converges in distribution to the standard normal distribution. In the Erdös-Rényi random graph, the variance of the average closure coefficient exhibits the same phase transition phenomenon as the average clustering coefficient.

math.ST

Asymptotic distributions of the average clustering coefficient and its variant

In network data analysis, summary statistics of a network can provide us with meaningful insight into the structure of the network. The average clustering coefficient is one of the most popular and widely used network statistics. In this paper, we investigate the asymptotic distributions of the average clustering coefficient and its variant of a heterogeneous Erdös-Rényi random graph. We show that the standardized average clustering coefficient converges in distribution to the standard normal distribution. Interestingly, the variance of the average clustering coefficient exhibits a phase transition phenomenon. The sum of weighted triangles is a variant of the average clustering coefficient. It is recently introduced to detect geometry in a network. We also derive the asymptotic distribution of the sum weighted triangles, which does not exhibit a phase transition phenomenon as the average clustering coefficient. This result signifies the difference between the two summary statistics.

math.ST

Asymptotic distribution of degree--based topological indices

Topological indices play a significant role in mathematical chemistry. Given a graph $\mathcal{G}$ with vertex set $\mathcal{V}=\{1,2,\dots,n\}$ and edge set $\mathcal{E}$, let $d_i$ be the degree of node $i$. The degree-based topological index is defined as $\mathcal{I}_n=$ $\sum_{\{i,j\}\in \mathcal{E}}f(d_i,d_j)$, where $f(x,y)$ is a symmetric function. In this paper, we investigate the asymptotic distribution of the degree-based topological indices of a heterogeneous Erdős-Rényi random graph. We show that after suitably centered and scaled, the topological indices converges in distribution to the standard normal distribution. Interestingly, we find that the general Randić index with $f(x,y)=(xy)^τ$ for a constant $τ$ exhibits a phase change at $τ=-\frac{1}{2}$.

math.CO

On the Randić index and its variants of network data

Summary statistics play an important role in network data analysis. They can provide us with meaningful insight into the structure of a network. The Randić index is one of the most popular network statistics that has been widely used for quantifying information of biological networks, chemical networks, pharmacologic networks, etc. A topic of current interest is to find bounds or limits of the Randić index and its variants. A number of bounds of the indices are available in literature. Recently, there are several attempts to study the limits of the indices in the Erdős-Rényi random graph by simulation. In this paper, we shall derive the limits of the Randić index and its variants of an inhomogeneous Erdős-Rényi random graph. Our results charaterize how network heterogeneity affects the indices and provide new insights about the Randić index and its variants. Finally we apply the indices to several real-world networks.

math.ST

Empirical likelihood test for community structure in networks

Network data, characterized by interconnected nodes and edges, is pervasive in various domains and has gained significant popularity in recent years. In network data analysis, testing the presence of community structure in a network is one of the important research tasks. Existing tests are mainly developed for unweighted networks. In this paper, we study the problem of testing the existence of community structure in general (either weighted or unweighted) networks. We propose two new tests: the Weighted Signed-Triangle (WST) test and the empirical likelihood (EL) test. Both tests can be applied to weighted or unweighted networks and outperform existing tests for small networks. The EL test may outperform the WST test for small networks.

stat.ME

On the Rényi index of random graphs

Networks (graphs) permeate scientific fields such as biology, social science, economics, etc. Empirical studies have shown that real-world networks are often heterogeneous, that is, the degrees of nodes do not concentrate on a number. Recently, the Rényi index was tentatively used to measure network heterogeneity. However, the validity of the Rényi index in network settings is not theoretically justified. In this paper, we study this problem. We derive the limit of the Rényi index of a heterogeneous Erdös-Rényi random graph and a power-law random graph, as well as the convergence rates. Our results show that the Erdös-Rényi random graph has asymptotic Rényi index zero and the power-law random graph (highly heterogeneous) has asymptotic Rényi index one. In addition, the limit of the Rényi index increases as the graph gets more heterogeneous. These results theoretically justify the Rényi index is a reasonable statistical measure of network heterogeneity. We also evaluate the finite-sample performance of the Rényi index by simulation.

math.ST

Information-theoretic Limits for Testing Community Structures in Weighted Networks

Community detection refers to the problem of clustering the nodes of a network into groups. Existing inferential methods for community structure mainly focus on unweighted (binary) networks. Many real-world networks are nonetheless weighted and a common practice is to dichotomize a weighted network to an unweighted one which is known to result in information loss. Literature on hypothesis testing in the latter situation is still missing. In this paper, we study the problem of testing the existence of community structure in weighted networks. Our contributions are threefold: (a). We use the (possibly infinite-dimensional) exponential family to model the weights and derive the sharp information-theoretic limit for the existence of consistent test. Within the limit, any test is inconsistent; and beyond the limit, we propose a useful consistent test. (b). Based on the information-theoretic limits, we provide the first formal way to quantify the loss of information incurred by dichotomizing weighted graphs into unweighted graphs in the context of hypothesis testing. (c). We propose several new and practically useful test statistics. Simulation study show that the proposed tests have good performance. Finally, we apply the proposed tests to an animal social network.

math.ST

Online Statistical Inference for Parameters Estimation with Linear-Equality Constraints

Stochastic gradient descent (SGD) and projected stochastic gradient descent (PSGD) are scalable algorithms to compute model parameters in unconstrained and constrained optimization problems. In comparison with SGD, PSGD forces its iterative values into the constrained parameter space via projection. From a statistical point of view, this paper studies the limiting distribution of PSGD-based estimate when the true parameters satisfy some linear-equality constraints. Our theoretical findings reveal the role of projection played in the uncertainty of the PSGD-based estimate. As a byproduct, we propose an online hypothesis testing procedure to test the linear-equality constraints. Simulation studies on synthetic data and an application to a real-world dataset confirm our theory.

stat.ML

Statistical Limits for Testing Correlation of Hypergraphs

In this paper, we consider the hypothesis testing of correlation between two $m$-uniform hypergraphs on $n$ unlabelled nodes. Under the null hypothesis, the hypergraphs are independent, while under the alternative hypothesis, the hyperdges have the same marginal distributions as in the null hypothesis but are correlated after some unknown node permutation. We focus on two scenarios: the hypergraphs are generated from the Gaussian-Wigner model and the dense Erdös-Rényi model. We derive the sharp information-theoretic testing threshold. Above the threshold, there exists a powerful test to distinguish the alternative hypothesis from the null hypothesis. Below the threshold, the alternative hypothesis and the null hypothesis are not distinguishable. The threshold involves $m$ and decreases as $m$ gets larger. This indicates testing correlation of hypergraphs ($m\geq3$) becomes easier than testing correlation of graphs ($m=2$)

math.ST

Community detection in censored hypergraph

Community detection refers to the problem of clustering the nodes of a network (either graph or hypergrah) into groups. Various algorithms are available for community detection and all these methods apply to uncensored networks. In practice, a network may has censored (or missing) values and it is shown that censored values have non-negligible effect on the structural properties of a network. In this paper, we study community detection in censored $m$-uniform hypergraph from information-theoretic point of view. We derive the information-theoretic threshold for exact recovery of the community structure. Besides, we propose a polynomial-time algorithm to exactly recover the community structure up to the threshold. The proposed algorithm consists of a spectral algorithm plus a refinement step. It is also interesting to study whether a single spectral algorithm without refinement achieves the threshold. To this end, we also explore the semi-definite relaxation algorithm and analyze its performance.

stat.ML