SearcharxivSearch

arXiv subjects

Allan Sly

Publications and source records attributed to Allan Sly.

At least 19 recordsLinked to original sources

Cutoff with an $O(1)$ window for Potts Glauber Dynamics on lattice at High Temperature

We prove cutoff with an $O(1)$ window for the continuous-time heat-bath Glauber dynamics of the ferromagnetic $q$-state Potts model on the discrete torus $\Lambda_n=(\mathbb Z/n\mathbb Z)^d$ at sufficiently high temperature. For every fixed $d\ge2$ and $q\ge3$, there exists $\beta_0=\beta_0(d,q)>0$ such that, for $0<\beta<\beta_0$, the Glauber dynamics of the Potts model on $\Lambda_n$ exhibits cutoff with optimal $O(1)$ window around \[ t_\star=t_\star^{(n)}:=\frac{1}{2\mathfrak{r}}\log |\Lambda_n|, \] where $\mathfrak{r}\in(0,1)$ is the exponential decay rate of the one-site magnetization. In particular, this determines the mixing time up to an additive $O(1)$. It is characterized by the point at which the macroscopic color-density bias from the monochromatic initial condition enters the scale of equilibrium fluctuations. Moreover, our proof shows that the monochromatic initial condition uniquely maximizes the color bias. This is the first implementation of information percolation to prove cutoff for a non-monotone spin system. In contrast with the Ising model, a direct implementation of information percolation does not yield matching upper and lower bounds for the Potts dynamics when $q\ge3$. We overcome this by developing an information-percolation framework for signed influences and combining it with Fourier bounds on signed convolution powers and geometric control of history diagrams.

math.PR

A Hierarchical Language Model with Predictable Scaling Laws and Provable Benefits of Reasoning

We introduce a family of synthetic languages with hierarchical structure -- generated by a broadcast process on trees -- for which the role of context length and reasoning in autoregressive generation can be analyzed precisely. At the heart of our analytic approach is an \emph{exact $k$-gram ansatz} in place of transformers with context length $k$, a substitution we then validate empirically. Using this ansatz we derive explicit asymptotic predictions for distributional statistics of the sequences produced by a trained model, instantiated in two settings. For the \emph{Ising broadcast process} (a soft-constrained language), we prove that the variance of the generated sum scales log-linearly in the context depth and its kurtosis converges to that of a Gaussian -- both deviating from the true language for any sublinear context. For the \emph{coloring broadcast process} (a hard-constrained language) in the freezing regime, bounded-context autoregression produces sequences that, with high probability, are inconsistent with \emph{any} valid coloring of the underlying tree. Together these results imply an $\Omega(n)$ lower bound on the context length required to faithfully sample length-$n$ sequences. In contrast, we prove that an autoregressive \emph{reasoning} model with only $\Theta(\log n)$ working memory can sample exactly from the true language -- an exponential improvement. We confirm both the lower-bound predictions and the reasoning-based upper bound empirically with transformers trained on the synthetic language; the trained models track our asymptotic predictions quantitatively across a wide range of context sizes.

cs.LG

Rapid phase ordering of Ising dynamics on $\mathbb Z^2$

We consider the phase ordering problem for the low-temperature Ising dynamics initialized from a biased and disordered initialization. Work of Fontes, Schonmann, Sidoravicius (2002) showed that at zero-temperature, Ising Glauber dynamics on $\mathbb Z^d$ for $d\ge 2$ initialized from i.i.d. spins on each vertex that are $+1$ with sufficiently large probability, absorbs into the all-plus configuration quickly. We prove that analogous behavior holds throughout the low-temperature regime of the Ising model in two dimensions. Namely, there exists $p_0 <1$ such that Ising Glauber dynamics initialized from i.i.d. spins that are $+1$ with probability $p>p_0$, run at any low temperature $\beta>\beta_c$ converges rapidly to the plus phase measure $\pi^+$. The result is proved using a spacetime multiscale coupling valid in any $d\ge 2$, that boosts a uniform-in-$\beta$ quasi-polynomial bound on the mixing time of Ising dynamics with plus boundary conditions, into rapid phase ordering from biased initializations with no boundary conditions.

math.PR

Rotationally invariant first passage percolation: Breaking the $n/\log n$ variance barrier

For first passage percolation (FPP) on Euclidean lattices $\mathbb{Z}^d$ with $d\ge 2$, it is expected that the variance of the first passage time between two points grows sublinearly in the distance with a universal exponent strictly smaller than $1$. Following Kesten's $O(n)$ upper bound (Ann. Appl. Probab., 1993) on the variance, Benjamini, Kalai and Schramm (Ann. Probab., 2003) used hypercontractivity to obtain an improvement of a factor of $\log n$ when passage times take two values with equal probability. This was later extended to more general classes of passage time distributions. However, unlike in exactly solvable planar models in last passage percolation where the variance is known to be $\Theta(n^{2/3})$, the best known upper bound for the variance of passage times has remained $O(n/\log n)$ in all non-trivial variants of FPP. For a class of rotationally invariant Riemannian FPP on the plane, we show that the variance is $O(n^{1-\varepsilon})$ for some $\varepsilon>0$. Our argument uses fluctuation estimates for passage times and geodesics derived in Basu, Sidoravicius and Sly (2023) together with a multi-scale argument to establish that the geodesic exhibits disorder chaos, i.e., upon resampling a small fraction of the underlying randomness, the updated geodesic has on average a small overlap with the original one; this, established at a large number of scales, leads to a polynomial improvement of the variance bound.

math.PR

Distribution of the magnetization of the critical Ising model on sparse random graphs

In this paper, we consider the Ising model on random $d$-regular graphs (with $d\ge3$) and Erd\"os-R\'enyi graphs $G(n,d/n)$ (with $d>1$) at the critical temperature. We prove that the \textit{magnetization}, i.e.\ the sum of the spins of a configuration, is typically of order $n^{3/4}$ and when multiplied by $n^{-3/4}$ converges in distribution to a non-trivial random variable, whose density we describe. In the regular graph case, the Small Subgraph Conditioning Method applies, and the limiting density is of the form $\frac1{Z}\exp(-C_d z^4)$. Surprisingly, in the Erd\"os-R\'enyi case, while the ratio of the second moment and first moment squared is bounded, the short cycle count is not enough to explain the fluctuations of the partition function restricted to a particular magnetization. We identify the additional source of randomness as path counts of slowly diverging length. This quantity is motivated by the heuristic that correlations between distant vertices are proportional to their local branching rate. Augmenting the Small Subgraph Conditioning Method with these path counts allows us to prove convergence of the magnetization to a non-deterministic limiting distribution. To our knowledge, the need to condition on graph observables beyond the cycle counts is a new phenomenon for spin systems. As further corollaries, we derive a polynomial lower bound on the mixing time of the stochastic Ising model on sparse random graphs at the critical temperature complementing recent upper bounds. Moreover, we establish the fluctuations of the free energy in the Erd\"os-R\'enyi case, answering a recent question of Coja-Oghlan et. al.

math.PR

Polynomial mixing of the critical Ising model on sparse Erdos-Renyi graphs

We consider the stochastic Ising model on sparse Erdos-Renyi graphs $G(n,d/n)$ with $d>1$ at the critical temperature $\beta_c=\tanh^{-1}(d^{-1})$ and prove that with high probability, the mixing time is at most polynomial in $n$. Our approach combines the recent stochastic localization framework of Chen and Eldan, which yields spectral gap bounds in the well-behaved bulk of the graph, together with classical results on the relaxation time of Glauber dynamics on trees to handle regions where we cannot apply the Chen-Eldan method directly because of atypically large local neighborhoods.

math.PR

Rapid phase ordering for Ising and Potts dynamics on random regular graphs

We consider the Ising, and more generally, $q$-state Potts Glauber dynamics on random $d$-regular graphs on $n$ vertices at low temperatures $\beta \gtrsim \frac{\log d}{d}$. The mixing time is exponential in $n$ due to a bottleneck between $q$ dominant phases consisting of configurations in which the majority of vertices are in the same state. We prove that for any $d\ge 7$, from biased initializations with $\epsilon_d n$ more vertices in state-$1$ than in other states, the Glauber dynamics quasi-equilibrates to the stationary distribution conditioned on having plurality in state-$1$ in optimal $O(\log n)$ time. Moreover, the requisite initial bias $\epsilon_d$ can be taken to zero as $d \to \infty$. Even for the $q=2$ Ising case, where the states are naturally identified with $\pm 1$, proving such a result requires a new approach in order to control negative information spread in spacetime despite the model being in low temperature and exhibiting strong local correlations. For this purpose, we introduce a coupled non-Markovian rigid dynamics for which a delicate temporal recursion on probability mass functions of minus spacetime cluster sizes establishes their subcriticality.

math.PR

Likelihood-Based Root State Reconstruction on a Tree: Sensitivity to Parameters and Applications

We consider a broadcasting problem on a tree where a binary digit (e.g., a spin or a nucleotide's purine/pyrimidine type) is propagated from the root to the leaves through symmetric noisy channels on the edges that randomly flip the state with edge-dependent probabilities. The goal of the reconstruction problem is to infer the root state given the observations at the leaves only. Specifically, we study the sensitivity of maximum likelihood estimation (MLE) to uncertainty in the edge parameters under this model, which is also known as the Cavender-Farris-Neyman (CFN) model. Our main result shows that when the true flip probabilities are sufficiently small, the posterior root mean (or magnetization of the root) under estimated parameters (within a constant factor) agrees with the root spin with high probability and deviates significantly from it with negligible probability. This provides theoretical justification for the practical use of MLE in ancestral sequence reconstruction in phylogenetics, where branch lengths (i.e., the edge parameters) must be estimated. As a separate application, we derive an approximation for the gradient of the population log-likelihood of the leaf states under the CFN model, with implications for branch length estimation via coordinate maximization.

math.PR

Polynomial Mixing of the critical Glauber Dynamics for the Ising Model

In this note, we prove that on any graph of maximal degree $d$ the mixing time of the Glauber Dynamics for the Ising Model at $\beta_c=\tanh^{-1}(\frac1{d-1})$, the uniqueness threshold on the infinite $d$-regular tree, is at most polynomial in $n$. The proof follows by a simple combination of new log-Sobolev bounds of Bauerschmidt and Dagallier, together with the tree of self avoiding walks construction of Weitz. While preparing this note we became aware that Chen, Chen, Yin and Zhang recently posted another proof of this result. We believe the simplicity of our argument is of independent interest.

math.PR

Weak recovery, hypothesis testing, and mutual information in stochastic block models and planted factor graphs

The stochastic block model is a canonical model of communities in random graphs. It was introduced in the social sciences and statistics as a model of communities, and in theoretical computer science as an average case model for graph partitioning problems under the name of the ``planted partition model.'' Given a sparse stochastic block model, the two standard inference tasks are: (i) Weak recovery: can we estimate the communities with non trivial overlap with the true communities? (ii) Detection/Hypothesis testing: can we distinguish if the sample was drawn from the block model or from a random graph with no community structure with probability tending to $1$ as the graph size tends to infinity? In this work, we show that for sparse stochastic block models, the two inference tasks are equivalent except at a critical point. That is, weak recovery is information theoretically possible if and only if detection is possible. We thus find a strong connection between these two notions of inference for the model. We further prove that when detection is impossible, an explicit hypothesis test based on low degree polynomials in the adjacency matrix of the observed graph achieves the optimal statistical power. This low degree test is efficient as opposed to the likelihood ratio test, which is not known to be efficient. Moreover, we prove that the asymptotic mutual information between the observed network and the community structure exhibits a phase transition at the weak recovery threshold. Our results are proven in much broader settings including the hypergraph stochastic block models and general planted factor graphs. In these settings we prove that the impossibility of weak recovery implies contiguity and provide a condition which guarantees the equivalence of weak recovery and detection.

math.PR

Infinite cycles in the interchange process in five dimensions

In the interchange process on a graph $G=(V,E)$, distinguished particles are placed on the vertices of $G$ with independent Poisson clocks on the edges. When the clock of an edge rings, the two particles on the two sides of the edge interchange. In this way, a random permutation $π_β:V\to V$ is formed for any time $β>0$. One of the main objects of study is the cycle structure of the random permutation and the emergence of long cycles. We prove the existence of infinite cycles in the interchange process on $\mathbb Z ^d$ for all dimensions $d\ge 5$ and all large $β$, establishing a conjecture of Bálint Tóth from 1993 in these dimensions. In our proof, we study a self-interacting random walk called the cyclic time random walk. Using a multiscale induction we prove that it is diffusive and can be coupled with Brownian motion. One of the key ideas in the proof is establishing a local escape property which shows that the walk will quickly escape when it is entangled in its history in complicated ways.

math.PR

Potts and random cluster measures on locally regular-tree-like graphs

Fixing $\beta \ge 0$ and an integer $q \ge 2$, consider the ferromagnetic $q$-Potts measures $\mu_n^{\beta,B}$ on finite graphs ${\sf G}_n$ on $n$ vertices, with external field strength $B \ge 0$ and the corresponding random cluster measures $\varphi^{q,\beta,B}_{n}$. Suppose that as $n \to \infty$ the uniformly sparse graphs ${\sf G}_n$ converge locally to an infinite $d$-regular tree ${\sf T}_{d}$, $d \ge 3$. We show that the convergence of the Potts free energy density to its Bethe replica symmetric prediction (which has been proved in case $d$ is even, or when $B=0$), yields the local weak convergence of $\varphi^{q,\beta,B}_n$ and $\mu_n^{\beta,B}$ to the corresponding free or wired random cluster measure, Potts measure, respectively, on ${\sf T}_{d}$. The choice of free versus wired limit is according to which has the larger Potts Bethe functional value, with mixtures of these two appearing {as limit points on} the critical line $\beta_c(q,B)$ where these two values of the Bethe functional coincide. For $B=0$ and $\beta>\beta_c$, we further establish a pure-state decomposition by showing that conditionally on the same dominant color $1 \le k \le q$, the $q$-Potts measures on such edge-expander graphs ${\sf G}_n$ converge locally to the $q$-Potts measure on ${\sf T}_{d}$ with a boundary wired at color $k$.

math.PR

Rotationally invariant first passage percolation: Concentration and scaling relations

For rotationally invariant first passage percolation (FPP) on the plane, we use a multi-scale argument to prove stretched exponential concentration of the first passage times at the scale of the standard deviation. Our results are proved under hypotheses which can be verified for many standard rotationally invariant models of first passage percolation, e.g. Riemannian FPP, Voronoi FPP and the Howard-Newman model. This is the first such tight concentration result known for any model that is not exactly solvable. As a consequence, we prove a version of the so called KPZ relation between the passage time fluctuations and the transversal fluctuations of geodesics as well as up to constant upper and lower bounds for the non-random fluctuations in these models. Similar results have previously been known conditionally under unproven hypotheses, but our results are the first ones that apply to some specific FPP models. Our arguments are expected to be useful in proving a number of other estimates which were hitherto only known conditionally or for exactly solvable models.

math.PR

Optimal reconstruction of general sparse stochastic block models

This paper is motivated by the reconstruction problem on the sparse stochastic block model. Mossel, et. al. proved that a reconstruction algorithm that recovers an optimal fraction of the communities in the symmetric, 2-community case. The main contribution of their proof is to show that when the signal to noise ratio is sufficiently large, in particular $λ^2d > C$, the reconstruction accuracy for a broadcast process on a tree with or without noise on the leaves is asymptotically the same. This paper will generalize their results, including the main step, to a general class of the sparse stochastic block model with any number of communities that are not necessarily symmetric, proving that an algorithm closely related to Belief Propagation recovers an optimal fraction of community labels.

math.PR

One-step replica symmetry breaking of random regular NAE-SAT II

Continuing our earlier work in \cite{nss20a}, we study the random regular k-NAE-SAT model in the condensation regime. In \cite{nss20a}, the 1RSB properties of the model were established with positive probability. In this paper, we improve the result to probability arbitrarily close to one. To do so, we introduce a new framework which is the synthesis of two approaches: the small subgraph conditioning and a variance decomposition technique using Doob martingales and discrete Fourier analysis. The main challenge is a delicate integration of the two methods to overcome the difficulty arising from applying the moment method to an unbounded state space.

math.PR

On a random model of forgetting

Georgiou, Katkov and Tsodyks considered the following random process. Let $x_1,x_2,\ldots $ be an infinite sequence of independent, identically distributed, uniform random points in $[0,1]$. Starting with $S=\{0\}$, the elements $x_k$ join $S$ one by one, in order. When an entering element is larger than the current minimum element of $S$, this minimum leaves $S$. Let $S(1,n)$ denote the content of $S$ after the first $n$ elements $x_k$ join. Simulations suggest that the size $|S(1,n)|$ of $S$ at time $n$ is typically close to $n/e$. Here we first give a rigorous proof that this is indeed the case, and that in fact the symmetric difference of $S(1,n)$ and the set $\{x_k\ge 1-1/e: 1 \leq k \leq n \}$ is of size at most $\tilde{O}(\sqrt n)$ with high probability. Our main result is a more accurate description of the process implying, in particular, that as $n$ tends to infinity $ n^{-1/2}\big( |S(1,n)|-n/e \big) $ converges to a normal random variable with variance $3e^{-2}-e^{-1}$. We further show that the dynamics of the symmetric difference of $S(1,n)$ and the set $\{x_k\ge 1-1/e: 1 \leq k \leq n \}$ converges with proper scaling to a three dimensional Bessel process.

math.PR

One-step replica symmetry breaking of random regular NAE-SAT I

In a broad class of sparse random constraint satisfaction problems(CSP), deep heuristics from statistical physics predict that there is a condensation phase transition before the satisfiability threshold, governed by one-step replica symmetry breaking(1RSB). In fact, in random regular k-NAE-SAT, which is one of such random CSPs, it was verified \cite{ssz22} that its free energy is well-defined and the explicit value follows the 1RSB prediction. However, for any model of sparse random CSP, it has been unknown whether the solution space indeed condenses on O(1) clusters according to the 1RSB prediction. In this paper, we give an affirmative answer to this question for the random regular k-NAE-SAT model. Namely, we prove that with probability bounded away from zero, most of the solutions lie inside a bounded number of solution clusters whose sizes are comparable to the scale of the free energy. Furthermore, we establish that the overlap between two independently drawn solutions concentrates precisely at two values. Our proof is based on a detailed moment analysis of a spin system, which has an infinite spin space that encodes the structure of solution clusters. We believe that our method is applicable to a broad range of random CSPs in the 1RSB universality class.

math.PR

The number of solutions for random regular NAE-SAT

Recent work has made substantial progress in understanding the transitions of random constraint satisfaction problems. In particular, for several of these models, the exact satisfiability threshold has been rigorously determined, confirming predictions of statistical physics. Here we revisit one of these models, random regular k-NAE-SAT: knowing the satisfiability threshold, it is natural to study, in the satisfiable regime, the number of solutions in a typical instance. We prove here that these solutions have a well-defined free energy (limiting exponential growth rate), with explicit value matching the one-step replica symmetry breaking prediction. The proof develops new techniques for analyzing a certain "survey propagation model" associated to this problem. We believe that these methods may be applicable in a wide class of related problems.

math.PR