SearcharxivSearch

arXiv subjects

David Aldous

Publications and source records attributed to David Aldous.

At least 19 recordsLinked to original sources

The Critical Beta-splitting Random Tree IV: Mellin analysis of Leaf Height

In the critical beta-splitting model of a random $n$-leaf rooted tree, clades are recursively split into sub-clades, and a clade of $m$ leaves is split into sub-clades containing $i$ and $m-i$ leaves with probabilities $\propto 1/(i(m-i))$. The height of a uniform random leaf can be represented as the absorption time of a certain {\em harmonic descent} Markov chain. Recent work on these heights $D_n$ and $L_n$ (corresponding to discrete or continuous versions of the tree) has led to quite sharp expressions for their asymptotic distributions, based on their Markov chain description. This article gives even sharper expressions, based on an $n \to \infty$ limit tree structure described via exchangeable random partitions in the style of Haas et al (2008). Within this structure, calculations of moments lead to expressions for Mellin transforms, and then via Mellin inversion we obtain sharp estimates for the expectation, variance, Normal approximation and large deviation behavior of $D_n$.

math.PR

The Critical Beta-splitting Random Tree: Heights and Related Results

In the critical beta-splitting model of a random $n$-leaf binary tree, leaf-sets are recursively split into subsets, and a set of $m$ leaves is split into subsets containing $i$ and $m-i$ leaves with probabilities proportional to $1/{i(m-i)}$. We study the continuous-time model in which the holding time before that split is exponential with rate $h_{m-1}$, the harmonic number. We (sharply) evaluate the first two moments of the time-height $D_n$ and of the edge-height $L_n$ of a uniform random leaf (that is, the length of the path from the root to the leaf), and prove the corresponding CLTs. We find the limiting value of the correlation between the heights of two random leaves of the same tree realization, and analyze the expected number of splits necessary for a set of $t$ leaves to partially or completely break away from each other. We give tail bounds for the time-height and the edge-height of the {\em tree}, that is the maximal leaf heights. We show that there is a limit distribution for the size of a uniform random subtree, and derive the asymptotics of the mean size. Our proofs are based on asymptotic analysis of the attendant (sum-type) recurrences. The essential idea is to replace such a recursive equality by a pair of recursive inequalities for which matching asymptotic solutions can be found, allowing one to bound, both ways, the elusive explicit solution of the recursive equality. This reliance on recursive inequalities necessitates usage of Laplace transforms rather than Fourier characteristic functions.

math.PR

Parking on the infinite binary tree

Let $(A_u : u \in \mathbb{B})$ be i.i.d.~non-negative integers that we interpret as car arrivals on the vertices of the full binary tree $ \mathbb{B}$. Each car tries to park on its arrival node, but if it is already occupied, it drives towards the root and parks on the first available spot. It is known that the parking process on $ \mathbb{B}$ exhibits a phase transition in the sense that either a finite number of cars do not manage to park in expectation (subcritical regime) or all vertices of the tree contain a car and infinitely many cars do not manage to park (supercritical regime). We characterize those regimes in terms of the law of $A$ in an explicit way. We also study in detail the critical regime as well as the phase transition which turns out to be "discontinuous".

math.PR

The Nearest Unvisited Vertex Walk on Random Graphs

We revisit an old minor topic in algorithms, the deterministic walk on a finite graph which always moves toward the nearest unvisited vertex until every vertex is visited. There is an elementary connection between this cover time and ball-covering (metric entropy) measures. For some familiar models of random graphs, this connection allows the order of magnitude of the cover time to be deduced from first passage percolation estimates. Establishing sharper results seems a challenging problem.

math.PR

The Life and Mathematical Legacy of Thomas M. Liggett

Thomas Milton Liggett was a world renowned UCLA probabilist, famous for his monograph Interacting Particle Systems. He passed away peacefully on May 12, 2020. This is a perspective article in memory of both Tom Liggett the person and Tom Liggett the mathematician.

math.PR

A Prediction Tournament Paradox

In a prediction tournament, contestants "forecast" by asserting a numerical probability for each of (say) 100 future real-world events. The scoring system is designed so that (regardless of the unknown true probabilities) more accurate forecasters will likely score better. This is true for one-on-one comparisons between contestants. But consider a realistic-size tournament with many contestants, with a range of accuracies. It may seem self-evident that the winner will likely be one of the most accurate forecasters. But, in the setting where the range extends to very accurate forecasters, simulations show this is mathematically false, within a somewhat plausible model. Even outside that setting the winner is less likely than intuition suggests to be one of the handful of best forecasters. Though implicit in recent technical papers, this paradox has apparently not been explicitly pointed out before, though is easily explained. It perhaps has implications for the ongoing IARPA-sponsored research programs involving forecasting.

math.ST

The optimal geometry of transportation networks

Motivated by the shape of transportation networks such as subways, we consider a distribution of points in the plane and ask for the network $G$ of given length $L$ that is optimal in a certain sense. In the general model, the optimality criterion is to minimize the average (over pairs of points chosen independently from the distribution) time to travel between the points, where a travel path consists of any line segments in the plane traversed at slow speed and any route within the subway network traversed at a faster speed. Of major interest is how the shape of the optimal network changes as $L$ increases. We first study the simplest variant of this problem where the optimization criterion is to minimize the average distance from a point to the network, and we provide some general arguments about the optimal networks. As a second variant we consider the optimal network that minimizes the average travel time to a central destination, and discuss both analytically and numerically some simple shapes such as the star network, the ring or combinations of both these elements. Finally, we discuss numerically the general model where the network minimizes the average time between all pairs of points. For this case, we propose a scaling form for the average time that we verify numerically. We also show that in the medium-length regime, as $L$ increases, resources go preferentially to radial branches and that there is a sharp transition at a value $L_c$ where a loop appears.

physics.soc-ph

To stay discovered: On tournament mean score sequences and the Bradley--Terry model

On being told that a piece of work he thought was his discovery had duplicated an earlier mathematician's work, Larry Shepp once replied "Yes, but when {\em I} discovered it, it {\em stayed} discovered". In this spirit we give discussion and probabilistic proofs of two related known results (Moon 1963, Joe 1988) on random tournaments which seem surprisingly unknown to modern probabilists. In particular our proof of Moon's theorem on mean score sequences seems more constructive than previous proofs. This provides a comparatively concrete introduction to a longstanding mystery, the lack of a canonical construction for a joint distribution in the representation theorem for convex order.

math.PR

Processes on Unimodular Random Networks

We investigate unimodular random networks. Our motivations include their characterization via reversibility of an associated random walk and their similarities to unimodular quasi-transitive graphs. We extend various theorems concerning random walks, percolation, spanning forests, and amenability from the known context of unimodular quasi-transitive graphs to the more general context of unimodular random networks. We give properties of a trace associated to unimodular random networks with applications to stochastic comparison of continuous-time random walk.

math.PR

Waves in a Spatial Queue: Stop-and-Go at Airport Security

We model a long queue of humans by a continuous-space model in which, when a customer moves forward, they stop a random distance behind the previous customer, but do not move at all if their distance behind the previous customer is below a threshold. The latter assumption leads to ``waves" of motion in which only some random number $W$ of customers move. We prove that $\Pr(W > k)$ decreases as order $k^{-1/2}$; in other words, for large $k$ the $k$'th customer moves on average only once every order $k^{1/2}$ service times. A more refined analysis relies on a non-obvious asymptotic relation to the coalescing Brownian motion process; we give a careful outline of such an analysis without attending to all the technical details.

math.PR

The Compulsive Gambler Process

In the compulsive gambler process there is a finite set of agents who meet pairwise at random times ($i$ and $j$ meet at times of a rate-$ν_{ij}$ Poisson process) and, upon meeting, play an instantaneous fair game in which one wins the other's money. We introduce this process and describe some of its basic properties. Some properties are rather obvious (martingale structure; comparison with Kingman coalescent) while others are more subtle (an "exchangeable over the money elements" property, and a construction reminiscent of the Donnelly-Kurtz look-down construction). Several directions for possible future research are described. One -- where agents meet neighbors in a sparse graph -- is studied here, and another -- a continuous-space extension called the {\em metric coalescent} -- is studied in Lanoue (2014).

math.PR

The Stretch - Length Tradeoff in Geometric Networks: Average Case and Worst Case Study

Consider a network linking the points of a rate-$1$ Poisson point process on the plane. Write $\Psi^{\mbox{ave}}(s)$ for the minimum possible mean length per unit area of such a network, subject to the constraint that the route-length between every pair of points is at most $s$ times the Euclidean distance. We give upper and lower bounds on the function $\Psi^{\mbox{ave}}(s)$, and on the analogous "worst-case" function $\Psi^{\mbox{worst}}(s)$ where the point configuration is arbitrary subject to average density one per unit area. Our bounds are numerically crude, but raise the question of whether there is an exponent $\alpha$ such that each function has $\Psi(s) \asymp (s-1)^{-\alpha}$ as $s \downarrow 1$.

math.PR

Interacting particle systems as stochastic social dynamics

The style of mathematical models known to probabilists as Interacting Particle Systems and exemplified by the Voter, Exclusion and Contact processes have found use in many academic disciplines. In many such disciplines the underlying conceptual picture is of a social network, where individuals meet pairwise and update their "state" (opinion, activity etc) in a way depending on the two previous states. This picture motivates a precise general setup we call Finite Markov Information Exchange (FMIE) processes. We briefly describe a few less familiar models (Averaging, Compulsive Gambler, Deference, Fashionista) suggested by the social network picture, as well as a few familiar ones.

math.ST

Another Conversation with Persi Diaconis

Persi Diaconis was born in New York on January 31, 1945. Upon receiving a Ph.D. from Harvard in 1974 he was appointed Assistant Professor at Stanford. Following periods as Professor at Harvard (1987-1997) and Cornell (1996-1998), he has been Professor in the Departments of Mathematics and Statistics at Stanford since 1998. He is a member of the National Academy of Sciences, a past President of the IMS and has received honorary doctorates from Chicago and four other universities. The following conversation took place at his office and at Aldous's home in early 2012.

stat.ME

Five Statistical Questions about the Tree of Life

Stochastic modeling of phylogenies raises five questions that have received varying levels of attention from quantitatively inclined biologists. 1) How large do we expect (from the model) the ration of maximum historical diversity to current diversity to be? 2) From a correct phylogeny of the extant species of a clade, what can we deduce about past speciation and extinction rates? 3) What proportion of extant species are in fact descendants of still-extant ancestral species, and how does this compare with predictions od models? 4) When one moves from trees on species to trees on sets of species (whether traditional higher order taxa or clades from PhyloCode), does one expect trees to become more unbiased as a purely logical consequence of tree structure, without signifying any real biological phenomenon? 5) How do we expect that fluctuation rates for counts of higher order taxa should compare with fluctuation rates for number of species? WE present a mathematician's view based on an oversimplified modeling framework in which all these questions can be studied coherently.

q-bio.PE

Fluctuations of Martingales and Winning Probabilities of Game Contestants

Within a contest there is some probability M_i(t) that contestant i will be the winner, given information available at time t, and M_i(t) must be a martingale in t. Assume continuous paths, to capture the idea that relevant information is acquired slowly. Provided each contestant's initial winning probability is at most b, one can easily calculate, without needing further model specification, the expectations of the random variables N_b = number of contestants whose winning probability ever exceeds b, and D_{ab} = total number of downcrossings of the martingales over an interval [a,b]. The distributions of N_b and D_{ab} do depend on further model details, and we study how concentrated or spread out the distributions can be. The extremal models for N_b correspond to two contrasting intuitively natural methods for determining a winner: progressively shorten a list of remaining candidates, or sequentially examine candidates to be declared winner or eliminated. We give less precise bounds on the variability of D_{ab}. We formalize the setting of infinitely many contestants each with infinitesimally small chance of winning, in which the explicit results are more elegant. A canonical process in this setting is the Wright-Fisher diffusion associated with an infinite population of initially distinct alleles; we show how this process fits our setting and raise the problem of finding the distributions of N_b and D_{ab} for this process.

math.PR

A Spatial Model of City Growth and Formation

We introduce a model in which city populations grow at rates proportional to the area of their "sphere of influence", where the influence of a city depends on its population (to power α) and distance from city (to power -β) and where new cities arise according to a certain random rule. A simple non-rigorous analysis of asymptotics indicates that for β> 2α$ the system exhibits "balanced growth" in which there are an increasing number of large cities, whose populations have the same order of magnitude, whereas for β< 2α$ the system exhibits "unbalanced growth" in which a few cities capture most of the total population. Conceptually the model is best regarded as a spatial analog of the combinatorial "Chinese restaurant process".

physics.soc-ph

Stochastic Models for Phylogenetic Trees on Higher-order Taxa

Simple stochastic models for phylogenetic trees on species have been well studied. But much paleontology data concerns time series or trees on higher-order taxa, and any broad picture of relationships between extant groups requires use of higher-order taxa. A coherent model for trees on (say) genera should involve both a species-level model and a model for the classification scheme by which species are assigned to genera. We present a general framework for such models, and describe three alternate classification schemes. Combining with the species-level model of Aldous-Popovic (2005), one gets models for higher-order trees, and we initiate analytic study of such models. In particular we derive formulas for the lifetime of genera, for the distribution of number of species per genus, and for the offspring structure of the tree on genera.

q-bio.PE