SearcharxivSearch

arXiv subjects

Qunqiang Feng

Publications and source records attributed to Qunqiang Feng.

10 recordsLinked to original sources

Optimal estimators and tests for reciprocal effects

The $p_1$ model plays a fundamental role in modeling directed networks, where the reciprocal effect parameter $ρ$ is of special interest in practice. However, due to nonlinear factors in this model, how to estimate $ρ$ efficiently is a long-standing open problem. We tackle the problem by the cycle count approach. The challenge is, due to the nonlinear factors in the model, for any given type of generalized cycles, the expected count is a complicated function of many parameters in the model, so it is unclear how to use cycle counts to estimate $ρ$. However, somewhat surprisingly, we discover that, among many types of generalized cycles with the same length, we can carefully pick a pair of them such that in the ratio between the expected cycle counts of the two types, the non-linear factors cancel out nicely with each other, and as a result, the ratio equals to $\mathrm{exp}(ρ)$ exactly. Therefore, though the expected count of cycles of any type is not tractable, the ratio between the expected cycle counts of a (carefully chosen) pair of generalized cycles may have an utterly simple form. We study to what extent such pairs exist, and use our discovery to derive both an estimate for $ρ$ and a testing procedure for testing $ρ= ρ_0$. In a setting where we allow a wide range of reciprocal effects and a wide variety of network sparsity and degree heterogeneity, we show that our estimator achieves the optimal rate and our test achieves the optimal phase transition. Technically, first, motivated by what we observe on real networks, we do not want to impose strong conditions on reciprocal effects, network sparsity, and degree heterogeneity. Second, our proposed statistic is a type of $U$-statistic, the analysis of which involves complex combinatorics and is error-prone. For these reasons, our analysis is long and delicate.

math.ST

Efficient Variance Estimation for the Polytomous Discrimination Index

Evaluating diagnostic accuracy for multi-category outcomes remains a significant challenge, primarily due to computational limitations in existing performance metrics. The Polytomous Discrimination Index (PDI) has emerged as an order-agnostic solution suitable for nominal classifications. However, its broader adoption has been constrained by the lack of efficient implementation, especially for its variance estimation, which typically would require computationally intensive bootstrapping procedures. In this work, we address this limitation by proposing a novel asymptotic variance estimator for the PDI. Our method integrates classical $U$-statistic theory with recent advances in combinatorics, offering a scalable and theoretically grounded alternative. To assess the performance of the proposed approach, we conduct extensive simulation studies and observe remarkable gain in computing time. We further apply our method to a real-world brain image analysis where deep neural networks are used as a diagnostic tool. We can efficiently report the accuracy of the neural networks with different depth specifications.

stat.CO

Power-law and log-periodic degree tails for a family of probability generating function equations arising in evolving networks

For a fixed integer $j\ge1$ and $0<p<1$, we study the probability generating function (pgf) equation \[ (1+2p)\,g(x)=2p\,x^{j}+g\bigl(x-px+px^{2}\bigr),\qquad 0\le x\le1 , \] which governs the limiting degree distribution $\{p_k\}$ of a family of evolving network models. The cases $j=1$ and $j=2$ are the treelike fast-growth model of Feng and Hu and the homogeneous evolving network of Feng, Li and Hu. We prove that for every $j$ the equation has a unique pgf solution, of mean $2j$, and we determine its coefficient tail exactly: \[ p_k=k^{-1-ρ}\,Ψ_j(\log_λk)+o\bigl(k^{-1-ρ}\bigr), \] where $λ=1+p$, $ρ=\log(1+2p)/\log(1+p)$ is independent of $j$, and $Ψ_j$ is continuous, strictly positive and $1$-periodic, with explicit Fourier coefficients. This resolves two conjectures of Feng and coauthors: (1) the power-law order $p_k=Θ(k^{-1-ρ})$ and (2) its refinement to the multiplicatively periodic form $p_k\simΨ_j(\log_λk)\,k^{-1-ρ}$. The periodic factor is genuinely non-constant for $p$ near $1$, and, for the two network models, for all $p$ outside a discrete set. Consequently, $p_k$ is asymptotic to no constant multiple of $k^{-1-ρ}$. Our method is a self-contained local analysis of the supercritical Galton-Watson process with offspring law $1+\mathrm{Bernoulli}(p)$, inspected at an independent geometric time. This time-changed process solves the equation observed by Feng and coauthors. The main results of this paper were obtained by the multi-agent system Eureka and have subsequently been verified by the authors.

math.PR

Subgraph counting estimation for the $β$-model in sparse networks

The $β$-model is popular for characterizing the commonly observed degree heterogeneity phenomenon in real-world networks. In this study, we develop a cycle counting approach to estimate $n$ node-specific parameters in the $β$-model for moderate or extremely sparse networks. Our proposed estimators, called \emph{Cycle Counting Ratio (CCR) Estimator}, are based on the log-ratios of two network cycle counting statistics with explicit expressions and therefore easy to compute. We focus on conditions to guarantee statistical properties of the single estimator for each node. Under the very weak conditions that $\max_t θ_t \to 0$ and $θ_t \|θ\|_1 \to \infty$, we show that the CCR estimator is consistent and achieves the minimax rate in terms of the mean squared error, which is the squared signal-to-noise ratio for $\hatβ_t$ up to a constant factor. Here, $\hatβ_t$ is the CCR estimator of the node-specific parameter $β_t$, $θ_t = \exp(β_t)$ and $θ=(θ_1, \ldots, θ_n)$. Even if the whole network density is close to the Erdős-Rényi lower bound $\log n/n$, the CCR estimator for the single parameter $β_t$ is still consistent as long as $θ_t \|θ\|_1 \to \infty$. To the best of our knowledge, this is the first time to derive the minimax rate and consistency result under such weak conditions. Under a slight stronger condition, we further establish its uniform consistency and asymptotic normality, whose asymptotic variance is $θ_t \|θ\|_1$. Numerical studies and an application to a sparse network data set demonstrate our theoretical findings.

stat.ME

Limit Laws for the Distance to Fréchet Means of Random Graphs

This paper investigates the Fréchet mean of the Erdős-Rényi random graph $G_{n,p}$ with respect to the Frobenius distance on graph Laplacians, a metric that captures global structural information beyond local edge flips. We first characterize the Fréchet mean set as consisting of quasi-regular graphs (i.e., graphs where all vertex degrees differ by at most one). We then analyze the asymptotic behavior of the Frobenius distance $F_n=d_{\mathrm{F}}(G_{n,p},R)$ as $n\to\infty$, where $R$ is any Fréchet mean. Closed-form expressions for the mean and variance of $F_n^2$ are derived, which are invariant to the choice of $R$. Leveraging these results, we establish several weak convergence laws for the Frobenius distance over all regimes of $p \in (0,1)$ as $n \to \infty$. Finally, under the scaling condition $n^2 p(1-p) \to \infty$ we prove the asymptotic normality of this distance, which exhibits a phase transition governed by the growth rate of $np(1-p)$. Our results reveal how metric selection fundamentally shapes Fréchet mean geometry in random graphs.

math.PR

Asymptotic normality for triangle counting in the sparse $β$-model

We study the number of triangles $T_n$ in the sparse $β$-model on $n$ vertices, a random graph model that captures degree heterogeneity in real-world networks. Using the norms of the heterogeneity parameter vector, we first determine the asymptotic mean and variance of $T_n$. Next, by applying the Malliavin-Stein method, we derive a non-asymptotic upper bound on the Kolmogorov distance between normalized $T_n$ and the standard normal distribution. Under an additional assumption on degree heterogeneity, we further prove the asymptotic normality for $T_n$, as $n\to\infty$.

math.PR

Triple-dyad ratio estimation for the $p_1$ model

Although the $p_1$ model was proposed 40 years ago, little progress has been made to address asymptotic theories in this model, that is, neither consistency of the maximum likelihood estimator (MLE) nor other parameter estimation with statistical guarantees is understood. This problem has been acknowledged as a long-standing open problem. To address it, we propose a novel parametric estimation method based on the ratios of the sum of a sequence of triple-dyad indicators to another one, where a triple-dyad indicator means the product of three dyad indicators. Our proposed estimators, called \emph{triple-dyad ratio estimator}, have explicit expressions and can be scaled to very large networks with millions of nodes. We establish the consistency and asymptotic normality of the triple-dyad ratio estimator when the number of nodes reaches infinity. Based on the asymptotic results, we develop a test statistic for evaluating whether is a reciprocity effect in directed networks. The estimators for the density and reciprocity parameters contain bias terms, where analytical bias correction formulas are proposed to make valid inference. Numerical studies demonstrate the findings of our theories and show that the estimator is comparable to the MLE in large networks.

stat.ME

The generalized Zagreb index for non-plane and plane recursive trees

The Zagreb index, which is defined as the sum of squares of degrees of the nodes of a tree, was studied in previous works by martingale techniques for random non-plane recursive trees and classes of random trees which are close to random plane recursive trees. These techniques are not easily amended to the generalized Zagreb index, which is defined similar but with squares replaced by higher powers. In this paper, we use the moment transfer approach to (i) obtain the first-order asymptotics of moments and to (ii) prove limit laws for the (suitable normalized) generalized Zagreb index for random non-plane and plane recursive trees; for the former, we show that for all higher powers the limit law is normal, for the latter, we show for cubes and fourth powers that its a non-normal law.

math.PR

Limit laws for the generalized Zagreb indices of random graphs

In this paper, we study the limiting behavior of the generalized Zagreb indices of the classical Erdős-Rényi (ER) random graph $G(n,p)$, as $n\to\infty$. For any integer $k\ge1$, we first give an expression for the $k$-th order generalized Zagreb index in terms of the number of star graphs of various sizes in any simple graph. The explicit formulas for the first two moments of the generalized Zagreb indices of an ER random graph are then obtained by this expression. Based on the asymptotic normality of the numbers of star graphs of various sizes, several joint limit laws are established for a finite number of generalized Zagreb indices with a phase transition for $p$ in different regimes. Finally, we provide a necessary and sufficient condition for any single generalized Zagreb index of $G(n,p)$ to be asymptotic normal.

math.PR

Average Jaccard Index of Random Graphs

The asymptotic behavior of the Jaccard index in $G(n,p)$, the classical Erdös-Rényi random graphs model, is studied in this paper, as $n$ goes to infinity. We first derive the asymptotic distribution of the Jaccard index of any pair of distinct vertices, as well as the first two moments of this index. Then the average of the Jaccard indices over all vertex pairs in $G(n,p)$ is shown to be asymptotically normal under an additional mild condition that $np\to\infty$ and $n^2(1-p)\to\infty$.

math.PR