SearcharxivSearch

arXiv subjects

Luca Avena

Publications and source records attributed to Luca Avena.

At least 19 recordsLinked to original sources

Renormalisation of Inhomogeneous Random Graphs

We consider inhomogeneous random graphs in which vertices are assigned i.i.d.\ random weights, pairs of distinct vertices are connected by an edge independently with a probability that is a bi-variate function of the weights of the vertices, and single vertices are connected to themselves by a self-loop independently with a probability that is a uni-variate function of the weight of the vertex. We apply a renormalisation transformation in which vertices are aggregated into groups of equal size according to a greedy algorithm, namely, distinct groups of aggregated vertices are connected by an aggregated edge if and only if there is at least one edge connecting two constituent vertices across the groups, while a group of aggregated vertices is connected to itself by an aggregated self-loop if and only if there is at least one self-loop at an internal vertex or one edge connecting a pair of distinct internal vertices. We analyse what happens when the renormalisation transformation is iterated. In particular, we show that, starting from appropriately scaled connection functions, the iterated renormalised graphs converge to a two-parameter family of random graphs, acting as an attractor in a universality class. We consider a light-tailed regime, for which the scaling limit is a homogeneous Erd\H{o}s--R\'enyi random graph, and a heavy-tailed regime, for which the scaling limit is an inhomogeneous random graph with stable infinite-mean random weights and an exponential disconnection function. Different scalings are needed for the two regimes. Which of the two regimes prevails depends on the choice of the connection functions and the choice of the law of the random weights.

math.PR

How reliable are LLMs when it comes to playing dice?

We investigate the probabilistic reasoning capabilities of large language models through a controlled benchmarking study on discrete probability problems. We constructed two datasets, respectively a set of standard exercises and a set of counterintuitive exercises, designed to trigger heuristic reasoning, and evaluated 8 state-of-the-art models, each tested with and without Chain-of-Thought prompting. Models achieve an average accuracy of 0.96 on standard problems but only 0.59 on counterintuitive ones. We further provide empirical evidence of token bias: performance drops by over 20% when canonical formulations are replaced by disguised variants. Embedding misleading suggestions in the prompt reduces performance by up to 34%, with no model proving immune. Taken together, the reported findings suggest that current LLMs are not yet genuine probabilistic reasoners, despite their success in advanced mathematical problems.

cs.CL

Counterintuitive problems in discrete probability

This manuscript contains a collection of counterintuitive problems in discrete probability, together with detailed solutions. The dataset was constructed as part of a broader research project investigating the capabilities of the latest-generation Large Language Models (LLMs) in solving discrete probability problems, in order to assess whether LLMs tend to make systematic reasoning errors associated with known cognitive biases. The problems collected here are specifically designed to challenge heuristic reasoning strategies that often lead to intuitively appealing but mathematically incorrect conclusions. The dataset combines several types of problems. Some are adapted from classical probabilistic paradoxes and cognitive-bias literature, while others originate from recreational mathematics sources or were developed by ourselves following similar principles. The primary purpose of this document is to provide a transparent and publicly accessible reference for the problems used in our experimental evaluation of language models, as well as providing detailed human-made solutions. At the same time, we believe that this collection may also prove useful for future research on probabilistic reasoning, cognitive biases, and the evaluation of reasoning capabilities in artificial intelligence systems.

math.PR

Network node immunization: improving Netshield algorithm through random rooted forests

We are interested in the so-called multiple-node immunization problem for complex networks under attack by a viral agent. It consists in identifying and removing a set of nodes of size $k$ in a graph to maximize the impeding of virus spread. A few approaches have been proposed in the literature based on numerical and theoretical insights on how classical models for virus spread evolve on graphs. Based on the analysis of these models, the maximal eigenvalue of the adjacency matrix of the graph has become a classical measure of how resilient the network is. Thus, a clear, well-explored approach for multiple-node immunization consists of identifying a set of $k$ nodes in such a way that the reduced network, obtained by removing these nodes, has a minimal largest eigenvalue. This spectral optimization problem turns out to be a computationally hard problem for which only greedy algorithms offer good solutions at efficient computational time. Among those, the so-called Netshield algorithm represents one of the reference choices. The latter is, in fact, a clearly defined algorithm aiming at optimizing a certain sub-modular functional, called shield-value, which approximates the original optimization problem. We propose here a novel procedure, based on random walk kernels and related random spanning forests, to build a new algorithm, referred to as K-shield, which enhances Netshield searching performance at the same computational complexity. We give theoretical insights behind this novel method, which could also be used for other optimization problems, and then test it via numerical showcase experiments on various benchmarks.

cs.SI

node2vec or triangle-biased random walks: stationarity, regularity & recurrence

The node2vec random walk is a non-Markovian random walk on the vertex set of a graph, widely used for network embedding and exploration. This random walk model is defined in terms of three parameters which control the probability of, respectively, backtracking moves, moves within triangles, and moves to the remaining neighboring nodes. From a mathematical standpoint, the node2vec random walk is a nontrivial generalization of the non-backtracking random walk and thus belongs to the class of second-order Markov chains. Despite its widespread use in applications, little is known about its long-run behavior. The goal of this paper is to begin exploring its fundamental properties on arbitrary graphs. To this aim, we show how lifting the node2vec random walk to the state spaces of directed edges and directed wedges yields two distinct Markovian representations which are key for its asymptotic analysis. Using these representations, we find mild sufficient conditions on the underlying finite or infinite graph to guarantee ergodicity, reversibility, recurrence and characterization of the invariant measure. As we discuss, the behavior of the node2vec random walk is drastically different compared to the non-backtracking random walk. While the latter simplifies on arbitrary graphs when using its natural edge Markovian representation thanks to bistochasticity, the former simplifies on regular graphs when using its natural wedge Markovian representation. Remarkably, this representation reveals that a graph is regular if and only if a certain weighted Eulerianity condition holds.

math.PR

Semicircle laws with combined variance for non-uniform Erd\H{o}s-R\'enyi hypergraphs

We consider Erd\H{o}s-R\'enyi-type random hypergraphs that are non-uniform, in the sense that hyperedges of different sizes may coexist, and inhomogeneous, in that connection probabilities may depend on the hyperedge size. All parameters are allowed to scale with the hypergraph size. We study the random adjacency matrix whose $(u,v)$-entry counts the number of hyperedges containing both vertices $u$ and $v$, and characterize its expected limiting spectral distribution in terms of the connection probabilities and the hyperedge sizes. We provide a Pastur-type condition, in the sense of Chatterjee (2005), under which the matrix can be Gaussianized, as well as a more restrictive but simpler sufficient condition in terms of the generalized average degree of the model. As a second main result, based on such a Gaussianization, we characterize the limiting spectral distributions under non-sparse conditions as semicircle laws with an explicit parametric variance. The latter can be expressed as a convex combination of the variances arising in the uniform cases, with coefficients determined by the trade-off between the different sources of inhomogeneity.

math.PR

A cohesive account on the ergodic behaviours and scaling limits of Random Walks in Cooling Random Environments

Transport in disordered media is a central theme in probability and statistical physics, where randomness in the underlying medium produces phenomena such as localization, anomalous scaling, and slow relaxation. A paradigmatic model for transport in disordered media is that of Random Walks in Random Environments (RWRE), which has been extensively studied since the 1970's and is by now well understood in one dimension. More recently, several works have explored perturbations of models of transport in disordered media aimed at interpolating between static disorder and fully homogenized dynamics. Random walks in cooling random environments (RWCRE), introduced in this context, constitute a key example: the environment is dynamically resampled at prescribed times and kept fixed in between, giving rise to a delicate ``quasi-ergodic'' structure in time allowing to interpolate between homogeneous random walks and classical RWRE. The purpose of this paper is twofold. A first goal is to offer an original survey on the main results for RWCRE in 1d: recurrence criteria, law of large numbers, large deviations, and fluctuation phenomena across different resampling regimes. Those results have been derived in a series of recent works and are complemented here with a number of new statements aiming at presenting a unified phenomenological picture. As a second goal, we try to extract a coherent conceptual picture that highlights structural mechanisms -- such as ergodic limits, persistence under perturbations and replacement principles -- that extend beyond the specific setting of RWCRE and are relevant for a broader class of disordered systems that can be perturbed by introducing independent resetting in the same fashion.

math.PR

Voter model on heterogeneous directed networks

We investigate the consensus dynamics of the voter model on large random graphs with heterogeneous and directed features, focusing in particular on networks with power-law degree distributions. By extending recent results on sparse directed graphs, we derive exact first-order asymptotics for the expected consensus time in directed configuration models with i.i.d. Pareto-distributed in- and out-degrees. For any tail exponent {\alpha}>0, we derive the mean consensus time scaling depending on the network size and a pre-factor that encodes detailed structural properties of the degree sequences. We give an explicit description of the pre factor in the directed setting. This extends and sharpens previous mean-field predictions from statistical physics, providing the first explicit consensus-time formula in the directed heavy-tailed setting. Through extensive simulations, we confirm the validity of our predictions across a wide range of heterogeneity regimes, including networks with infinite variance and infinite mean degree distribution. We further explore the interplay between network topology and voter dynamics, highlighting how degree fluctuations and maximal degrees shape the consensus landscape. Complementing the asymptotic analysis, we provide numerical evidence for the emergence of Wright-Fisher diffusive behavior in both directed and undirected ensembles under suitable mixing conditions, and demonstrate the breakdown of this approximation in the in the infinite mean regime.

math.PR

The voter model on random regular graphs with random rewiring

We consider the voter model with binary opinions on a random regular graph with $n$ vertices of degree $d \geq 3$, subject to a rewiring dynamics in which pairs of edges are rewired, i.e., broken into four half-edges and subsequently reconnected at random. A parameter $\nu \in (0,\infty)$ regulates the frequency at which the rewirings take place, in such a way that any given edge is rewired exponentially at a rate $\nu$ in the limit as $n\to\infty$. We show that, under the joint law of the random rewiring dynamics and the random opinion dynamics, the fraction of vertices with either one of the two opinions converges on time scale $n$ to the Fisher-Wright diffusion with an explicit diffusion constant $\vartheta_{d,\nu}$ in the limit as $n\to\infty$. In particular, we identify $\vartheta_{d,\nu}$ in terms of a continued-fraction expansion and analyse its dependence on $d$ and $\nu$. A key role in our analysis is played by the set of discordant edges, which constitutes the boundary between the sets of vertices carrying the two opinions.

math.PR

Discordant edges for the voter model on regular random graphs

We consider the two-opinion voter model on a regular random graph with n vertices and degree $d \geq 3$. It is known that consensus is reached on time scale n and that on this time scale the volume of the set of vertices with one opinion evolves as a Fisher-Wright diffusion. We are interested in the evolution of the number of discordant edges (i.e., edges linking vertices with different opinions), which can be thought as the perimeter of the set of vertices with one opinion, and is the key observable capturing how consensus is reached. We show that if initially the two opinions are drawn independently from a Bernoulli distribution with parameter $u \in (0, 1)$, then on time scale 1 the fraction of discordant edges decreases and stabilises to a value that depends on d and u, and is related to the meeting time of two random walks on an infinite tree of degree d starting from two neighbouring vertices. Moreover, we show that on time scale n the fraction of discordant edges moves away from the constant plateau and converges to zero in an exponential fashion. Our proofs exploit the classical dual system of coalescing random walks and use ideas from Cooper et al. (2010) built on the so-called First Visit Time Lemma. We further introduce a novel technique to derive concentration properties from weak-dependence of coalescing random walks on moderate time scales.

math.PR

Mixing of fast random walks on dynamic random permutations

We analyse the mixing profile of a random walk on a dynamic random permutation, focusing on the regime where the walk evolves much faster than the permutation. Two types of dynamics generated by random transpositions are considered: one allows for coagulation of permutation cycles only, the other allows for both coagulation and fragmentation. We show that for both types, after scaling time by the length of the permutation and letting this length tend to infinity, the total variation distance between the current distribution and the uniform distribution converges to a limit process that drops down in a single jump. This jump is similar to a one-sided cut-off, occurs after a random time whose law we identify, and goes from the value 1 to a value that is a strictly decreasing and deterministic function of the time of the jump, related to the size of the largest component in Erd\H{o}s-R\'enyi random graphs. After the jump, the total variation distance follows this function down to 0.

math.PR

Limiting Spectra of inhomogeneous random graphs

We consider sparse inhomogeneous Erdős-Rényi random graph ensembles where edges are connected independently with probability $p_{ij}$. We assume that $p_{ij}= \varepsilon_N f(w_i, w_j)$ where $(w_i)_{i\ge 1}$ is a sequence of deterministic weights, $f$ is a bounded function and $N\varepsilon_N\to λ\in (0,\infty)$. We characterise the limiting moments in terms of graph homomorphisms and also classify the contributing partitions. We present an analytic way to determine the Stieltjes transform of the limiting measure. The convergence of the empirical distribution function follows from the theory of local weak convergence in many examples but we do not rely on this theory and exploit combinatorial and analytic techniques to derive some interesting properties of the limit. We extend the methods of Khorunzhy et al. (2004) and show that a fixed point equation determines the limiting measure. The limiting measure crucially depends on $λ$ and it is known that in the homogeneous case, if $λ\to\infty$, the measure converges weakly to the semicircular law (Jung and Lee (2018)). We extend this result of interpolating between the sparse and dense regimes to the inhomogeneous setting and show that as $λ\to \infty$, the measure converges weakly to a measure which is known as the operator-valued semicircular law.

math.PR

Meeting, coalescence and consensus time on random directed graphs

We consider Markovian dynamics on a typical realization of the so-called Directed Configuration Model (DCM), that is, a random directed graph with prescribed in- and out-degrees. In this random geometry, we study the meeting time of two random walks on a typical realization of the graph starting at stationarity, the coalescence time for a system of coalescent random walks, and the consensus time of the voter model. Indeed, it is known that the latter three quantities are related to each other when the underlying sequence of graphs satisfies certain mean field conditions. Such conditions can be summarized by requiring a fast mixing time of the random walk and some anti-concentration of its stationary distribution: properties that a typical random directed graph is known to have under natural assumptions on the degree sequence. In this paper we show that, for a typical large graph from the DCM ensemble, the distribution of the meeting time is well-approximated by an exponential random variable and we provide the first-order approximation of its expectation, showing that the latter is linear in the size of the graph, and the preconstant depends on some easy statistics of the degree sequence. As a byproduct, we are able to analyze the effect of the degree sequence in changing the meeting, coalescence and consensus time. Our approach follows the classical idea of converting meeting into hitting times of a proper collapsed chain, which we control by the so-called First Visit Time Lemma.

math.PR

Loop-erased partitioning of a network: monotonicities & analysis of cycle-free graphs

We consider random partitions of the vertex set of a given finite graph that can be sampled by means of loop-erased random walks stopped at a random exponential time of parameter $q>0$. The related random blocks tend to cluster nodes visited by the random walk on time scale $1/q$. This random partitioning is induced by a measure of rooted spanning forest of the graph, which generalizes the classical uniform spanning tree measure and which can be obtained as a zero-limit of FK-percolation with an external cemetery state. Some general properties of this rooted forest measure and related determinantal observables, along with a number of applications in data analysis have been recently explored. We are here mainly interested in the structure the emergent partitioning, referred to as loop-erased partitioning, as the scale parameter $q$ varies. We first present two general results shedding light on subtle monoticity properties in $q$ of these rooted forest and loop-erased partitioning measures. The first theorem characterizes monotone events in $q$ by deriving a Russo-like formula. Our second general result concerns two-point correlations defined by the probability that two vertices do not belong to the same block of the partitioning. It states that, on undirected graphs, these correlation functions are increasing in $q$. We then explore other types of results aiming at understanding the emerging asymptotic clusters on simple growing graph models, as $q$ scales with the graph size. Some first results in this direction have been investigated in the recent [arXiv:1906.03858] on dense geometries. Here we look at very sparse graphs. We offer a detailed analysis of the resulting partitioning on line segments and we look at trees and other tree-like geometries, without and with implanted modular structures. For the latter, we characterize the asymptotic detection of these implanted modules.

math.PR

Inhomogeneous random graphs with infinite-mean fitness variables

We consider an inhomogeneous Erd\H{o}s-R\'enyi random graph ensemble with exponentially decaying random disconnection probabilities determined by an i.i.d. field of variables with heavy tails and infinite mean associated to the vertices of the graph. This model was recently investigated in the physics literature in Garuccio et al. (2020) as a scale-invariant random graph within the context of network renormalization. From a mathematical perspective, the model fits in the class of scale-free inhomogeneous random graphs whose asymptotic geometrical features have been recently attracting interest. While for this type of graphs several results are known when the underlying vertex variables have finite mean and variance, here instead we consider the case of one-sided stable variables with necessarily infinite mean. To simplify our analysis, we assume that the variables are sampled from a Pareto distribution with parameter $\alpha\in(0,1)$. We start by characterizing the asymptotic distributions of the typical degrees and some related observables. In particular, we show that the degree of a vertex converges in distribution, after proper scaling, to a mixed Poisson law. We then show that correlations among degrees of different vertices are asymptotically non-vanishing, but at the same time a form of asymptotic tail independence is found when looking at the behavior of the joint Laplace transform around zero. Moreover, we present some findings concerning the asymptotic density of wedges and triangles and show a cross-over for the existence of dust (i.e. disconnected vertices).

math.PR

Gaussian, stable, tempered stable and mixed limit laws for random walks in cooling random environments

Random Walks in Cooling Random Environments (RWCRE) is a model of random walks in dynamic random environments where the entire environment is resampled along a fixed sequence of times, called the "cooling sequence," and is kept fixed in between those times. This model interpolates between that of a homogenous random walk, where the environment is reset at every step, and Random Walks in (static) Random Environments (RWRE), where the environment is never resampled. In this work we focus on the limiting distributions of one-dimensional RWCRE in the regime where the fluctuations of the corresponding (static) RWRE is given by a $s$-stable random variable with $s\in(1,2)$. In this regime, due to the two extreme cases (resampling every step and never resampling, respectively), a crossover from Gaussian to stable limits for sufficiently regular cooling sequence was previously conjectured. Our first result answers affirmatively this conjecture by making clear critical exponent, norming sequences and limiting laws associated with the crossover which demonstrates a change from Gaussian to $s$-stable limits, passing at criticality through a certain generalized tempered stable distribution. We then explore the resulting RWCRE scaling limits for general cooling sequences. On the one hand, we offer sets of operative sufficient conditions that guarantee asymptotic emergence of either Gaussian, $s$-stable or generalized tempered distributions from a certain class. On the other hand, we give explicit examples and describe how to construct irregular cooling sequences for which the corresponding limit law is characterized by mixtures of the three above mentioned laws. To obtain these results, we need and derive a number of refined asymptotic results for the static RWRE with $s\in(1,2)$ which may be of independent interest.

math.PR

Laws of large numbers for weighted sums of independent random variables: a game of mass

We consider weighted sums of independent random variables regulated by an increment sequence. We provide operative conditions that ensure strong law of large numbers for such sums to hold in both the centered and non-centered case. The existing criteria for the strong law are either implicit or assume some sufficient decay for the sequence of coefficients. In our set up we allow for arbitrary sequence of coefficients, possibly random, provided the random variables regulated by such increments satisfy some mild concentration conditions. In the non-centered case, convergence can be translated into the behavior of a deterministic sequence and it becomes a game of mass provided the expectation of the random variables is a function of the increments. We show how different limiting scenarios can emerge by identifying several classes of increments, for which concrete examples will be offered.

math.PR

Linking the mixing times of random walks on static and dynamic random graphs

This paper considers non-backtracking random walks on random graphs generated according to the configuration model. The quantity of interest is the scaling of the mixing time of the random walk as the number of vertices of the random graph tends to infinity. Subject to mild general conditions, we link two mixing times: one for a static version of the random graph, the other for a class of dynamic versions of the random graph in which the edges are randomly rewired but the degrees are preserved. The link is provided by the probability that the random walk has not yet stepped along a previously rewired edge. We use this link to compute the scaling of the mixing time for three specific classes of random rewirings. Depending on the speed and the range of the rewiring relative to the current location of the random walk, the mixing time may exhibit no cut-off, one-sided cut-off or two-sided cut-off, a trichotomy that was also found in earlier work. Interestingly, for a class of dynamics that are `mesoscopic', i.e., non-local and non-global, we find new behaviour with six subregimes. Proofs are built on a new and flexible coupling scheme, in combination with sharp estimates on the degrees encountered by the random walk in the static and the dynamic version of the random graph. Some of these estimates require sharp control on possible short-cuts in the graph between the edges that are traversed by the random walk.

math.PR