SearcharxivSearch

arXiv subjects

David J. Aldous

Publications and source records attributed to David J. Aldous.

At least 19 recordsLinked to original sources

The Critical Beta-splitting Random Tree II: Overview and Open Problems

In the critical beta-splitting model of a random $n$-leaf rooted tree, clades are recursively (from the root) split into sub-clades, and a clade of $m$ leaves is split into sub-clades containing $i$ and $m-i$ leaves with probabilities $\propto 1/(i(m-i))$. Study of structure theory and explicit quantitative aspects of this model (in discrete or continuous versions) is an active research topic. For many results there are different proofs, probabilistic or analytic, so the model provides a testbed for a ``compare and contrast" discussion of techniques. This article provides an overview of results proved in the sequence of similarly-titled articles I, III, IV and related articles. We mostly do not repeat proofs given elsewhere: instead we seek to paint a ``Big Picture" via graphics and heuristics, and emphasize open problems. Our discussion is centered around three categories of results. (i) There is a CLT for leaf heights, and the analytic proofs can be extended to provide surprisingly precise analysis of other height-related aspects. (ii) There is an explicit description of the limit {\em fringe distribution} relative to a random leaf, whose graphical representation is essentially the format of the cladogram representation of biological phylogenies. (iii) There is a canonical embedding of the discrete model into a continuous-time model, that is a random tree CTCS(n) on $n$ leaves with real-valued edge lengths, and this model turns out more convenient to study. The family (CTCS(n), n \ge 2) is consistent under a ``delete random leaf and prune" operation. That leads to an explicit inductive construction of (CTCS(n), n \ge 2) as $n$ increases, and then to a limit structure CTCS($\infty$) formalized via exchangeable partitions. Many open problems remain, in particular to elucidate a relation between CTCS($\infty$) and the $β(2,1)$ coalescent.

math.PR

The Critical Beta-splitting Random Tree III: The exchangeable partition representation and the fringe tree

In the critical beta-splitting model of a random $n$-leaf rooted tree, clades are recursively split into sub-clades, and a clade of $m$ leaves is split into sub-clades containing $i$ and $m-i$ leaves with probabilities $\propto 1/(i(m-i))$. Study of structure theory and explicit quantitative aspects of the model is an active research topic. It turns out that many results have several different proofs, and detailed studies of analytic proofs are given elsdewhere (via analysis of recursions and via Mellin transforms). This article describes two core probabilistic methods for studying $n \to \infty$ asymptotics of the basic finite-$n$-leaf models. (i) There is a canonical embedding into a continuous-time model, that is a random tree CTCS(n) on $n$ leaves with real-valued edge lengths, and this model turns out to be more convenient to study. The family (CTCS(n), $n \ge 2)$ is consistent under a ``delete random leaf and prune" operation. That leads to an explicit inductive construction (the {\em growth algorithm}) of (CTCS(n), $n \ge 2)$ as $n$ increases, and then to a limit structure CTCS$(\infty)$ which can be formalized via exchangeable partitions, in some ways analogous to the Brownian continuum random tree. (ii) There is an explicit description of the limit fringe distribution relative to a random leaf, whose graphical representation is essentially the format of the cladogram representation of biological phylogenies.

math.PR

The Harmonic Descent Chain

The decreasing Markov chain on \{1,2,3, \ldots\} with transition probabilities $p(j,j-i) \propto 1/i$ arises as a key component of the analysis of the beta-splitting random tree model. We give a direct and almost self-contained "probability" treatment of its occupation probabilities, as a counterpart to a more sophisticated but perhaps opaque derivation using a limit continuum tree structure and Mellin transforms.

math.PR

Markov chains and mappings of distributions on compact spaces

Consider a compact metric space $S$ and a pair $(j,k)$ with $k \ge 2$ and $1 \le j \le k$. For any probability distribution $θ\in P(S)$, define a Markov chain on $S$ by: from state $s$, take $k$ i.i.d. ($θ$) samples, and jump to the $j$'th closest. Such a chain converges in distribution to a unique stationary distribution, say $π_{j,k}(θ)$. So this defines a mapping $π_{j,k}: P(S) \to P(S)$. What happens when we iterate this mapping? In particular, what are the fixed points of this mapping? We present a few rigorous results, to complement our extensive simulation study elsewhere.

math.PR

Markov chains and mappings of distributions on compact spaces II: Numerics and Conjectures

Consider a compact metric space $S$ and a pair $(j,k)$ with $k \ge 2$ and $1 \le j \le k$. For any probability distribution $θ\in P(S)$, define a Markov chain on $S$ by: from state $s$, take $k$ i.i.d. ($θ$) samples, and jump to the $j$'th closest. Such a chain converges in distribution to a unique stationary distribution, say $π_{j,k}(θ)$. This defines a mapping $π_{j,k}: P(S) \to P(S)$. What happens when we iterate this mapping? In particular, what are the fixed points of this mapping? A few results are proved in a companion article; this article, not intended for formal publication, records numerical studies and conjectures.

math.PR

The distance problem on measured metric spaces

What distributions arise as the distribution of the distance between two typical points in some measured metric space? This seems to be a surprisingly subtle problem. We conjecture that every distribution with a density function whose support contains $0$ does arise in this way, and give some partial results in that direction.

math.PR

Gambling under unknown probabilities as a proxy for real world decisions under uncertainty

We give elementary examples within a framework for studying decisions under uncertainty where probabilities are only roughly known. The framework, in gambling terms, is that the size of a bet is proportional to the gambler's perceived advantage based on their perceived probability, and their accuracy in estimating true probabilities is measured by mean squared-error. Within this framework one can study the cost of estimation errors, and seek to formalize the ``obvious" notion that in competitive interactions between agents whose actions depend on their perceived probabilities, those who are more accurate at estimating probabilities will generally be more successful than those who are less accurate.

math.PR

On the Largest Common Subtree of Random Leaf-Labeled Binary Trees

The size of the largest common subtree (maximum agreement subtree) of two independent uniform random binary trees on $n$ leaves is known to be between orders $n^{1/8}$ and $n^{1/2}$. By a construction based on recursive splitting and analyzable by standard "stochastic fragmentation" methods, we improve the lower bound to order $n^β$ for $β= \frac{\sqrt{3} - 1}{2} = 0.366$. Improving the upper bound remains a challenging problem.

math.PR

A Real-World Markov Chain arising in Recreational Volleyball

Card shuffling models have provided simple motivating examples for the mathematical theory of mixing times for Markov chains. As a complement, we introduce a more intricate realistic model of a certain observable real-world scheme for mixing human players onto teams. We quantify numerically the effectiveness of this mixing scheme over the 7 or 8 steps performed in practice. We give a combinatorial proof of the non-trivial fact that the chain is indeed irreducible.

math.PR

Route Lengths in Invariant Spatial Tree Networks

Is there a constant $r_0$ such that, in any invariant tree network linking rate-$1$ Poisson points in the plane, the mean within-network distance between points at Euclidean distance $r$ is infinite for $r > r_0$? We prove a slightly weaker result. This is a continuum analog of a result of Benjamini et al (2001) on invariant spanning trees of the integer lattice.

math.PR

Covering a compact space by fixed-radius or growing random balls

Simple random coverage models, well studied in Euclidean space, can also be defined on a general compact metric space. By analogy with the geometric models, and with the discrete coupon collector's problem and with cover times for finite Markov chains, one expects a "weak concentration" bound for the distribution of the cover time to hold under minimal assumptions. We give two such results, one for random fixed-radius balls and the other for sequentially arriving randomly-centered and deterministically growing balls. Each is in fact a simple application of a different more general bound, the former concerning coverage by i.i.d. random sets with arbitrary distribution, and the latter concerning hitting times for Markov chains with a strong monotonicity property. The growth model seems generally more tractable, and we record some basic results and open problems for that model.

math.PR

Random partitions of the plane via Poissonian coloring, and a self-similar process of coalescing planar partitions

Plant differently colored points in the plane, then let random points ("Poisson rain") fall, and give each new point the color of the nearest existing point. Previous investigation and simulations strongly suggest that the colored regions converge (in some sense) to a random partition of the plane. We prove a weak version of this, showing that normalized empirical measures converge to Lebesgue measures on a random partition into measurable sets. Topological properties remain an open problem. In the course of the proof, which heavily exploits time-reversals, we encounter a novel self-similar process of coalescing planar partitions. In this process, sets $A(z)$ in the partition are associated with Poisson random points $z$, and the dynamics are as follows. Points are deleted randomly at rate $1$, when $z$ is deleted, its set $A(z)$ is adjoined to the set $A(z^\prime)$ of the nearest other point $z^\prime$.

math.PR

Weak Concentration for First Passage Percolation Times on Graphs and General Increasing Set-valued Processes

A simple lemma bounds $\mathrm{s.d.}(T)/\mathbb{E} T$ for hitting times $T$ in Markov chains with a certain strong monotonicity property. We show how this lemma may be applied to several increasing set-valued processes. Our main result concerns a model of first passage percolation on a finite graph, where the traversal times of edges are independent Exponentials with arbitrary rates. Consider the percolation time $X$ between two arbitrary vertices. We prove that $\mathrm{s.d.}(X)/\mathbb{E} X$ is small if and only if $Ξ/\mathbb{E} X$ is small, where $Ξ$ is the maximal edge-traversal time in the percolation path attaining $X$.

math.PR

Entropy of Some Models of Sparse Random Graphs With Vertex-Names

Consider the setting of sparse graphs on N vertices, where the vertices have distinct "names", which are strings of length O(log N) from a fixed finite alphabet. For many natural probability models, the entropy grows as cN log N for some model-dependent rate constant c. The mathematical content of this paper is the (often easy) calculation of c for a variety of models, in particular for various standard random graph models adapted to this setting. Our broader purpose is to publicize this particular setting as a natural setting for future theoretical study of data compression for graphs, and (more speculatively) for discussion of unorganized versus organized complexity.

math.PR

Scale-Invariant Random Spatial Networks

Real-world road networks have an approximate scale-invariance property; can one devise mathematical models of random networks whose distributions are {\em exactly} invariant under Euclidean scaling? This requires working in the continuum plane. We introduce an axiomatization of a class of processes we call {\em scale-invariant random spatial networks}, whose primitives are routes between each pair of points in the plane. We prove that one concrete model, based on minimum-time routes in a binary hierarchy of roads with different speed limits, satisfies the axioms, and note informally that two other constructions (based on Poisson line processes and on dynamic proximity graphs) are expected also to satisfy the axioms. We initiate study of structure theory and summary statistics for general processes in this class.

math.PR

Connected Spatial Networks over Random Points and a Route-Length Statistic

We review mathematically tractable models for connected networks on random points in the plane, emphasizing the class of proximity graphs which deserves to be better known to applied probabilists and statisticians. We introduce and motivate a particular statistic $R$ measuring shortness of routes in a network. We illustrate, via Monte Carlo in part, the trade-off between normalized network length and $R$ in a one-parameter family of proximity graphs. How close this family comes to the optimal trade-off over all possible networks remains an intriguing open question. The paper is a write-up of a talk developed by the first author during 2007--2009.

math.PR

When Knowing Early Matters: Gossip, Percolation and Nash Equilibria

Continually arriving information is communicated through a network of $n$ agents, with the value of information to the $j$'th recipient being a decreasing function of $j/n$, and communication costs paid by recipient. Regardless of details of network and communication costs, the social optimum policy is to communicate arbitrarily slowly. But selfish agent behavior leads to Nash equilibria which (in the $n \to \infty$ limit) may be efficient (Nash payoff $=$ social optimum payoff) or wasteful ($0 < $ Nash payoff $<$ social optimum payoff) or totally wasteful (Nash payoff $=0$). We study the cases of the complete network (constant communication costs between all agents), the grid with only nearest-neighbor communication, and the grid with communication cost a function of distance. The main technical tool is analysis of the associated first passage percolation process or SI epidemic (representing spread of one item of information) and in particular its "window width", the time interval during which most agents learn the item. Many arguments are just outlined, not intended as complete rigorous proofs.

math.PR