SearcharxivSearch

arXiv subjects

Shuyang Gong

Publications and source records attributed to Shuyang Gong.

12 recordsLinked to original sources

The Conclave Process

We introduce a stochastic model for the papal conclave in which $n$ cardinals vote repeatedly among themselves until one cardinal receives all the votes. In each round, the probability that a cardinal votes for a given candidate is proportional to the $α$-th power of that candidate's vote count in the preceding round. For $α=1$, the model reduces to the Wright-Fisher model and is dual to Kingman's n-coalescent. We reveal a sharp transition in the absorption time $\mathcal{T}$ at $α=1$. It was known that when $α=1$, $\mathcal{T}$ is typically of order $n$. We prove that for $α>1$, it drops to order $\textit{loglog n.}$ In contrast, for $α<1$, $\mathcal{T}$ is typically at least $\exp(Ω(n))$. We also prove a sharp phase transition in the identity of the winner when $α>1$. For every positive integer $k$, if $2^{1/k}<α<2^{1/(k-1)}$ (where we write $2^{1/0} = +\infty$), with probability tending to 1 as $n\to\infty$, the eventual winner is the unique leader after round $k$. These results show that reinforced voting processes reach consensus remarkably quickly even for large electorates.

math.PR

The Umeyama algorithm for matching correlated Gaussian geometric models in the low-dimensional regime

Motivated by the problem of matching two correlated random geometric graphs, we study the problem of matching two Gaussian geometric models correlated through a latent node permutation. Specifically, given an unknown permutation $π^*$ on $\{1,\ldots,n\}$ and given $n$ i.i.d. pairs of correlated Gaussian vectors $\{X_{π^*(i)},Y_i\}$ in $\mathbb{R}^d$ with noise parameter $σ$, we consider two types of (correlated) weighted complete graphs with edge weights given by $A_{i,j}=\langle X_i,X_j \rangle$, $B_{i,j}=\langle Y_i,Y_j \rangle$. The goal is to recover the hidden vertex correspondence $π^*$ based on the observed matrices $A$ and $B$. For the low-dimensional regime where $d=O(\log n)$, Wang, Wu, Xu, and Yolou [WWXY22+] established the information thresholds for exact and almost exact recovery in matching correlated Gaussian geometric models. They also conducted numerical experiments for the classical Umeyama algorithm. In our work, we prove that this algorithm achieves exact recovery of $π^*$ when the noise parameter $σ=o(d^{-3}n^{-2/d})$, and almost exact recovery when $σ=o(d^{-3}n^{-1/d})$. Our results approach the information thresholds up to a $\operatorname{poly}(d)$ factor in the low-dimensional regime.

math.ST

Detection and Reconstruction of a Random Hypergraph from Noisy Graph Projection

For a $d$-uniform random hypergraph on $n$ vertices in which hyperedges are included i.i.d.\ so that the average degree in the hypergraph is $n^{δ+o(1)}$, the projection of such a hypergraph is a graph on the same $n$ vertices where an edge connects two vertices if and only if they belong to a same hyperedge. In this work, we study the inference problem where the observation is a \emph{noisy} version of the graph projection where each edge in the projection is kept with probability $p=n^{-1+α+o(1)}$ and each edge not in the projection is added with probability $q=n^{-1+β+o(1)}$. For all constant $d$, we establish sharp thresholds for both detection (distinguishing the noisy projection from an Erdős-Rényi random graph with edge density $q$) and reconstruction (estimating the original hypergraph). Notably, our results reveal a \emph{detection-reconstruction gap} phenomenon in this problem. Our work also answers a problem raised in \cite{BGPY25+}.

math.ST

Stable algorithms cannot reliably find isolated perceptron solutions

We study the binary perceptron, a random constraint satisfaction problem that asks to find a Boolean vector in the intersection of independently chosen random halfspaces. A striking feature of this model is that at every positive constraint density, it is expected that a $1-o_N(1)$ fraction of solutions are \emph{strongly isolated}, i.e. separated from all others by Hamming distance $Ω(N)$. At the same time, efficient algorithms are known to find solutions at certain positive constraint densities. This raises a natural question: can any isolated solution be algorithmically visible? We answer this in the negative: no algorithm whose output is stable under a tiny Gaussian resampling of the disorder can \emph{reliably} locate isolated solutions. We show that any stable algorithm has success probability at most $\frac{3\sqrt{17}-9}{4}+o_N(1)\leq 0.84233$. Furthermore, every stable algorithm that finds a solution with probability $1-o_N(1)$ finds an isolated solution with probability $o_N(1)$. The class of stable algorithms we consider includes degree-$D$ polynomials up to $D\leq o(N/\log N)$; under the low-degree heuristic \cite{hopkins2018statistical}, this suggests that locating strongly isolated solutions requires running time $\exp(\widetildeΘ(N))$. Our proof does not use the overlap gap property. Instead, we show via Pitt's correlation inequality that after a random perturbation of the disorder, the number of solutions located close to a pre-existing isolated solution cannot concentrate at $1$.

cs.CC

Fundamental Limits of Community Detection in Contextual Multi-Layer Stochastic Block Models

We consider the problem of community detection from the joint observation of a high-dimensional covariate matrix and $L$ sparse networks, all encoding noisy, partial information about the latent community labels of $n$ subjects. In the asymptotic regime where the networks have constant average degree and the number of features $p$ grows proportionally with $n$, we derive a sharp threshold under which detecting and estimating the subject labels is possible. Our results extend the work of \cite{MN23} to the constant-degree regime with noisy measurements, and also resolve a conjecture in \cite{YLS24+} when the number of networks is a constant. Our information-theoretic lower bound is obtained via a novel comparison inequality between Bernoulli and Gaussian moments, as well as a statistical variant of the ``recovery to chi-square divergence reduction'' argument inspired by \cite{DHSS25}. On the algorithmic side, we design efficient algorithms based on counting decorated cycles and decorated paths and prove that they achieve the sharp threshold for both detection and weak recovery. In particular, our results show that there is no statistical-computational gap in this setting.

math.ST

Detecting Correlation Efficiently in Stochastic Block Models: Breaking Otter's Threshold in the Entire Supercritical Regime

Consider a pair of sparse correlated stochastic block models $\mathcal S(n,\tfracλ{n},ε;s)$ subsampled from a common parent stochastic block model with two symmetric communities, average degree $λ=O(1)$, divergence parameter $ε\in (0,1)$ and subsampling probability $s$. For all $ε\in(0,1)$ and $Δ>0$, we construct a statistic based on the combination of two low-degree polynomials and show that there exists a sufficiently small constant $δ=δ(ε,λ,Δ)>0$ such that if $ε^2 λs>1+Δ$ and $s>\sqrtα-δ$ where $α\approx 0.338$ is Otter's constant, this statistic can distinguish this model and a pair of independent stochastic block models $\mathcal S(n,\tfrac{λs}{n},ε)$ with probability $1-o(1)$. We also provide an efficient algorithm that approximates this statistic in polynomial time. The crux of our statistic's construction lies in a carefully curated family of multigraphs called \emph{decorated trees}, which enables effective aggregation of the community signal and graph correlation by leveraging the counts of the same decorated tree while suppressing the undesirable correlations among counts of different decorated trees. We believe such construction may be of independent interest.

cs.DS

Finding a dense submatrix of a random matrix. Sharp bounds for online algorithms

We consider the problem of finding a dense submatrix of a matrix with i.i.d. Gaussian entries, where density is measured by average value. This problem arose from practical applications in biology and social sciences \cites{madeira-survey,shabalin2009finding} and is known to exhibit a computation-to-optimization gap between the optimal value and best values achievable by existing polynomial time algorithms. In this paper we consider the class of online algorithms, which includes the best known algorithm for this problem, and derive a tight approximation factor ${4\over 3\sqrt{2}}$ for this class. The result is established using a simple implementation of recently developed Branching-Overlap-Gap-Property \cite{huang2025tight}. We further extend our results to $(\mathbb R^n)^{\otimes p}$ tensors with i.i.d. Gaussian entries, for which the approximation factor is proven to be ${2\sqrt{p}/(1+p)}$.

math.PR

A computational transition for detecting correlated stochastic block models by low-degree polynomials

Detection of correlation in a pair of random graphs is a fundamental statistical and computational problem that has been extensively studied in recent years. In this work, we consider a pair of correlated (sparse) stochastic block models $\mathcal{S}(n,\tfracλ{n};k,ε;s)$ that are subsampled from a common parent stochastic block model $\mathcal S(n,\tfracλ{n};k,ε)$ with $k=O(1)$ symmetric communities, average degree $λ=O(1)$, divergence parameter $ε$, and subsampling probability $s$. For the detection problem of distinguishing this model from a pair of independent Erdős-Rényi graphs with the same edge density $\mathcal{G}(n,\tfrac{λs}{n})$, we focus on tests based on \emph{low-degree polynomials} of the entries of the adjacency matrices, and we determine the threshold that separates the easy and hard regimes. More precisely, we show that this class of tests can distinguish these two models if and only if $s> \min \{ \sqrtα, \frac{1}{λε^2} \}$, where $α\approx 0.338$ is the Otter's constant and $\frac{1}{λε^2}$ is the Kesten-Stigum threshold. Combining a reduction argument in \cite{Li25+}, our hardness result also implies low-degree hardness for partial recovery and detection (to independent block models) when $s< \min \{ \sqrtα, \frac{1}{λε^2} \}$. Finally, our proof of low-degree hardness is based on a conditional variant of the low-degree likelihood calculation.

math.PR

A Proof of The Changepoint Detection Threshold Conjecture in Preferential Attachment Models

We investigate the problem of detecting and estimating a changepoint in the attachment function of a network evolving according to a preferential attachment model on $n$ vertices, using only a single final snapshot of the network. Bet et al.~\cite{bet2023detecting} show that a simple test based on thresholding the number of vertices with minimum degrees can detect the changepoint when the change occurs at time $n-Ω(\sqrt{n})$. They further make the striking conjecture that detection becomes impossible for any test if the change occurs at time $n-o(\sqrt{n}).$ Kaddouri et al.~\cite{kaddouri2024impossibility} make a step forward by proving the detection is impossible if the change occurs at time $n-o(n^{1/3}).$ In this paper, we resolve the conjecture affirmatively, proving that detection is indeed impossible if the change occurs at time $n-o(\sqrt{n}).$ Furthermore, we establish that estimating the changepoint with an error smaller than $o(\sqrt{n})$ is also impossible, thereby confirming that the estimator proposed in Bhamidi et al.~\cite{bhamidi2018change} is order-optimal.

math.PR

Asymptotic diameter of preferential attachment model

We study the asymptotic diameter of the preferential attachment model $\operatorname{PA}\!_n^{(m,δ)}$ with parameters $m \ge 2$ and $δ> 0$. Building on the recent work \cite{VZ25}, we prove that the diameter of $G_n \sim \operatorname{PA}\!_n^{(m,δ)}$ is $(1+o(1))\log_νn$ with high probability, where $ν$ is the exponential growth rate of the local weak limit of $G_n$. Our result confirms the conjecture in \cite{VZ25} and closes the remaining gap in understanding the asymptotic diameter of preferential attachment graphs with general parameters $m \ge 1$ and $δ>-m$. Our proof follows a general recipe that relates the diameter of a random graph to its typical distance, which we expect to have applicability in a broader range of models.

math.PR

The Algorithmic Phase Transition of Random Graph Alignment Problem

We study the graph alignment problem over two independent Erdős-Rényi graphs on $n$ vertices, with edge density $p$ falling into two regimes separated by the critical window around $p_c=\sqrt{\log n/n}$. Our result reveals an algorithmic phase transition for this random optimization problem: polynomial-time approximation schemes exist in the sparse regime, while statistical-computational gap emerges in the dense regime. Additionally, we establish a sharp transition on the performance of online algorithms for this problem when $p$ lies in the dense regime, resulting in a $\sqrt{8/9}$ multiplicative constant factor gap between achievable and optimal solutions.

math.PR

A polynomial-time approximation scheme for the maximal overlap of two independent Erdős-Rényi graphs

For two independent Erdős-Rényi graphs $\mathbf G(n,p)$, we study the maximal overlap (i.e., the number of common edges) of these two graphs over all possible vertex correspondence. We present a polynomial-time algorithm which finds a vertex correspondence whose overlap approximates the maximal overlap up to a multiplicative factor that is arbitrarily close to 1. As a by-product, we prove that the maximal overlap is asymptotically $\frac{n}{2α-1}$ for $p=n^{-α}$ with some constant $α\in (1/2,1)$.

math.PR