Searcharxiv⌕ Search

arXiv subjects

Jim Pitman

Publications and source records attributed to Jim Pitman.

At least 37 records · Page 2Linked to original sources

A representation of exchangeable hierarchies by sampling from real trees

A hierarchy on a set $S$, also called a total partition of $S$, is a collection $\mathcal{H}$ of subsets of $S$ such that $S \in \mathcal{H}$, each singleton subset of $S$ belongs to $\mathcal{H}$, and if $A, B \in \mathcal{H}$ then $A \cap B$ equals either $A$ or $B$ or $\varnothing$. Every exchangeable random hierarchy of positive integers has the same distribution as a random hierarchy $\mathcal{H}$ associated as follows with a random real tree $\mathcal{T}$ equipped with root element $0$ and a random probability distribution $p$ on the Borel subsets of $\mathcal{T}$: given $(\mathcal{T},p)$, let $t_1,t_2, ...$ be independent and identically distributed according to $p$, and let $\mathcal{H}$ comprise all singleton subsets of $\mathbb{N}$, and every subset of the form $\{j: t_j \in F_x\}$ as $x$ ranges over $\mathcal{T}$, where $F_x$ is the fringe subtree of $\mathcal{T}$ rooted at $x$. There is also the alternative characterization: every exchangeable random hierarchy of positive integers has the same distribution as a random hierarchy $\mathcal{H}$ derived as follows from a random hierarchy $\mathscr{H}$ on $[0,1]$ and a family $(U_j)$ of IID uniform [0,1] random variables independent of $\mathscr{H}$: let $\mathcal{H}$ comprise all sets of the form $\{j: U_j \in B\}$ as $B$ ranges over the members of $\mathscr{H}$.

math.PR↗

Ordered and size-biased frequencies in GEM and Gibbs models for species sampling

We describe the distribution of frequencies ordered by sample values in a random sample of size $n$ from the two parameter GEM$(α,θ)$ random discrete distribution on the positive integers. These frequencies are a $($size$-α)$-biased random permutation of the sample frequencies in either ranked order, or in the order of appearance of values in the sampling process. This generalizes a well known identity in distribution due to Donnelly and Tavaré (1986) for $α= 0$ to the case $0 \le α< 1$. This description extends to sampling from Gibbs$(α)$ frequencies obtained by suitable conditioning of the GEM$(α,θ)$ model, and yields a value-ordered version of the Chinese Restaurant construction of GEM$(α,θ)$ and Gibbs$(α)$ frequencies in the more usual size-biased order of their appearance. The proofs are based on a general construction of a finite sample $(X_1,\dots,X_n)$ from any random frequencies in size-biased order from the associated exchangeable random partition $Π_\infty$ of $\mathbb{N}$ which they generate.

math.PR↗

The spans in Brownian motion

For $d \in \{1,2,3\}$, let $(B^d_t;~ t \geq 0)$ be a $d$-dimensional standard Brownian motion. We study the $d$-Brownian span set $Span(d):=\{t-s;~ B^d_s=B^d_t~\mbox{for some}~0 \leq s \leq t\}$. We prove that almost surely the random set $Span(d)$ is $σ$-compact and dense in $\mathbb{R}_{+}$. In addition, we show that $Span(1)=\mathbb{R}_{+}$ almost surely; the Lebesgue measure of $Span(2)$ is $0$ almost surely and its Hausdorff dimension is $1$ almost surely; and the Hausdorff dimension of $Span(3)$ is $\frac{1}{2}$ almost surely. We also list a number of conjectures and open problems.

math.PR↗

An ergodic theorem for partially exchangeable random partitions

We consider shifts $Π_{n,m}$ of a partially exchangeable random partition $Π_\infty$ of $\mathbb{N}$ obtained by restricting $Π_\infty$ to $\{n+1,n+2,\dots, n+m\}$ and then subtracting $n$ from each element to get a partition of $[m]:= \{1, \ldots, m \}$. We show that for each fixed $m$ the distribution of $Π_{n,m}$ converges to the distribution of the restriction to $[m]$ of the exchangeable random partition of $\mathbb{N}$ with the same ranked frequencies as $Π_\infty$. As a consequence, the partially exchangeable random partition $Π_\infty$ is exchangeable if and only if $Π_\infty$ is stationary in the sense that for each fixed $m$ the distribution of $Π_{n,m}$ on partitions of $[m]$ is the same for all $n$. We also describe the evolution of the frequencies of a partially exchangeable random partition under the shift transformation. For an exchangeable random partition with proper frequencies, the time reversal of this evolution is the heaps process studied by Donnelly and others.

math.PR↗

Extremes and gaps in sampling from a GEM random discrete distribution

We show that in a sample of size $n$ from a GEM$(0,θ)$ random discrete distribution, the gaps $G_{i:n}:= X_{n-i+1:n} - X_{n-i:n}$ between order statistics $X_{1:n} \le \cdots \le X_{n:n}$ of the sample, with the convention $G_{n:n} := X_{1:n} - 1$, are distributed like the first $n$ terms of an infinite sequence of independent geometric$(i/(i+θ))$ variables $G_i$. This extends a known result for the minimum $X_{1:n}$ to other gaps in the range of the sample, and implies that the maximum $X_{n:n}$ has the distribution of $1 + \sum_{i=1}^n G_i$, hence the known result that $X_{n:n}$ grows like $θ\log(n)$ as $n\to\infty$, with an asymptotically normal distribution. Other consequences include most known formulas for the exact distributions of GEM$(0,θ)$ sampling statistics, including the Ewens and Donnelly--Tavaré sampling formulas. For the two-parameter GEM$(α,θ)$ distribution we show that the maximal value grows like a random multiple of $n^{α/(1-α)}$ and find the limit distribution of the multiplier.

math.PR↗

Successive maxima of samples from a GEM distribution

We show that the maximal value in a size $n$ sample from GEM$(θ)$ distribution is distributed as a sum of independent geometric random variables. This implies that the maximal value grows as $θ\log(n)$ as $n\to\infty$. For the two-parametric GEM$(α,θ)$ distribution we show that the maximal value grows as a random factor of $n^{α/(1-α)}$ and find the limiting distribution.

math.PR↗

A direct approach to the stable distributions

The explicit form for the characteristic function of a stable distribution on the line is derived analytically by solving the associated functional equation and applying theory of regular variation, without appeal to the general Lévy-Khintchine integral representation of infinitely divisible distributions.

math.PR↗

Regenerative tree growth: Markovian embedding of fragmenters, bifurcators, and bead splitting processes

Some, but not all processes of the form $M_t=\exp(-ξ_t)$ for a pure-jump subordinator $ξ$ with Laplace exponent $Φ$ arise as residual mass processes of particle 1 (tagged particle) in Bertoin's partition-valued exchangeable fragmentation processes. We introduce the notion of a Markovian embedding of $M=(M_t,t\ge 0)$ in a fragmentation process, and we show that for each $Φ$, there is a unique (in distribution) binary fragmentation process in which $M$ has a Markovian embedding. The identification of the Laplace exponent $Φ^*$ of its tagged particle process $M^*$ gives rise to a symmetrisation operation $Φ\mapstoΦ^*$, which we investigate in a general study of pairs $(M,M^*)$ that coincide up to a random time and then evolve independently. We call $M$ a fragmenter and $(M,M^*)$ a bifurcator. For $α>0$, we equip the interval $R_1=[0,\int_0^{\infty}M_t^α\,dt]$ with a purely atomic probability measure $μ_1$, which captures the jump sizes of $M$ suitably placed on $R_1$. We study binary tree growth processes that in the $n$th step sample an atom (``bead'') from $μ_n$ and build $(R_{n+1},μ_{n+1})$ by replacing the atom by a rescaled independent copy of $(R_1,μ_1)$ that we tie to the position of the atom. We show that any such bead splitting process $((R_n,μ_n),n\ge1)$ converges almost surely to an $α$-self-similar continuum random tree of Haas and Miermont, in the Gromov-Hausdorff-Prohorov sense. This generalises Aldous's line-breaking construction of the Brownian continuum random tree.

math.PR↗

Size-biased permutation of a finite sequence with independent and identically distributed terms

This paper focuses on the size-biased permutation of $n$ independent and identically distributed (i.i.d.) positive random variables. This is a finite dimensional analogue of the size-biased permutation of ranked jumps of a subordinator studied in Perman-Pitman-Yor (PPY) [Probab. Theory Related Fields 92 (1992) 21-39], as well as a special form of induced order statistics [Bull. Inst. Internat. Statist. 45 (1973) 295-300; Ann. Statist. 2 (1974) 1034-1039]. This intersection grants us different tools for deriving distributional properties. Their comparisons lead to new results, as well as simpler proofs of existing ones. Our main contribution, Theorem 25 in Section 6, describes the asymptotic distribution of the last few terms in a finite i.i.d. size-biased permutation via a Poisson coupling with its few smallest order statistics.

math.PR↗

Patterns in random walks and Brownian motion

We ask if it is possible to find some particular continuous paths of unit length in linear Brownian motion. Beginning with a discrete version of the problem, we derive the asymptotics of the expected waiting time for several interesting patterns. These suggest corresponding results on the existence/non-existence of continuous paths embedded in Brownian motion. With further effort we are able to prove some of these existence and non-existence results by various stochastic analysis arguments. A list of open problems is presented.

math.PR↗

Beta-gamma tail asymptotics

We compute the tail asymptotics of the product of a beta random variable and a generalized gamma random variable which are independent and have general parameters. A special case of these asymptotics were proved and used in a recent work of Bubeck, Mossel, and Rácz in order to determine the tail asymptotics of the maximum degree of the preferential attachment tree. The proof presented here is simpler and highlights why these asymptotics hold.

math.PR↗

Random Dirichlet series arising from records

We study the distributions of the random Dirichlet series with parameters $(s, β)$ defined by $$ S=\sum_{n=1}^{\infty}\frac{I_n}{n^s}, $$ where $(I_n)$ is a sequence of independent Bernoulli random variables, $I_n$ taking value $1$ with probability $1/n^β$ and value $0$ otherwise. Random series of this type are motivated by the record indicator sequences which have been studied in extreme value theory in statistics. We show that when $s>0$ and $0< β\le 1$ with $s+β>1$ the distribution of $S$ has a density; otherwise it is purely atomic or not defined because of divergence. In particular, in the case when $s>0$ and $β=1$, we prove that for every $0 1$ it is unbounded. In the case when $s>0$ and $0<β<1$ with $s+β>1$, the density is smooth. To show the absolute continuity, we obtain estimates of the Fourier transforms, employing van der Corput's method to deal with number-theoretic problems. We also give further regularity results of the densities, and present an example of non atomic singular distribution which is induced by the series restricted to the primes.

math.PR↗

The Slepian zero set, and Brownian bridge embedded in Brownian motion by a spacetime shift

This paper is concerned with various aspects of the Slepian process $(B_{t+1} - B_t, t \ge 0)$ derived from a one-dimensional Brownian motion $(B_t, t \ge 0 )$. In particular, we offer an analysis of the local structure of the Slepian zero set $\{t : B_{t+1} = B_t \}$, including a path decomposition of the Slepian process for $0 \le t \le 1$. We also establish the existence of a random time $T$ such that $T$ falls in the the Slepian zero set almost surely and the process $(B_{T+u} - B_T, 0 \le u \le 1)$ is standard Brownian bridge.

math.PR↗

The Vervaat transform of Brownian bridges and Brownian motion

For a continuous function $f \in \mathcal{C}([0,1])$, define the Vervaat transform $V(f)(t):=f(τ(f)+t \mod1)+f(1)1_{\{t+τ(f) \geq 1\}}-f(τ(f))$, where $τ(f)$ corresponds to the first time at which the minimum of $f$ is attained. Motivated by recent study of quantile transforms of random walks and Brownian motion, we investigate the Vervaat transform of Brownian motion and Brownian bridges with arbitrary endpoints. When the two endpoints of the bridge are not the same, the Vervaat transform is not Markovian. We describe its distribution by path decomposition and study its semi-martingale property. The same study is done for the Vervaat transform of unconditioned Brownian motion, the expectation and variance of which are also derived.

math.PR↗

Martingale marginals do not always determine convergence

Baez-Duarte (1971) and Gilat (1972) gave examples of martingales that converge in probability (and hence in distribution) but not almost surely. Here such a martingale is constructed with uniformly bounded increments, and a construction is provided of two martingales with the same marginals, one of which converges almost surely, while the other does not converge in probability.

math.PR↗

On Vervaat transform of Brownian bridges and Brownian motion

For a continuous function $f \in \mathcal{C}([0,1])$, define the Vervaat transform $V(f)(t):=f(τ(f)+t \mod1)+f(1)1_{\{t+τ(f) \geq 1\}}-f(τ(f))$, where $τ(f)$ corresponds to the first time at which the minimum of $f$ is attained. Motivated by recent study of quantile transforms for random walks and Brownian motion, we study the Vervaat transform of Brownian motion and Brownian bridges with arbitary endpoints. When the two endpoints of the bridge are not the same, the Vervaat transform is not Markovian. We describe its distribution by path decompositions and study its semimartingale properties. The expectation and variance of the Vervaat transform of Brownian motion are also derived.

math.PR↗

Cluster and Feature Modeling from Combinatorial Stochastic Processes

One of the focal points of the modern literature on Bayesian nonparametrics has been the problem of clustering, or partitioning, where each data point is modeled as being associated with one and only one of some collection of groups called clusters or partition blocks. Underlying these Bayesian nonparametric models are a set of interrelated stochastic processes, most notably the Dirichlet process and the Chinese restaurant process. In this paper we provide a formal development of an analogous problem, called feature modeling, for associating data points with arbitrary nonnegative integer numbers of groups, now called features or topics. We review the existing combinatorial stochastic process representations for the clustering problem and develop analogous representations for the feature modeling problem. These representations include the beta process and the Indian buffet process as well as new representations that provide insight into the connections between these processes. We thereby bring the same level of completeness to the treatment of Bayesian nonparametric feature modeling that has previously been achieved for Bayesian nonparametric clustering.

math.ST↗

Regenerative tree growth: structural results and convergence

We introduce regenerative tree growth processes as consistent families of random trees with n labelled leaves, n>=1, with a regenerative property at branch points. This framework includes growth processes for exchangeably labelled Markov branching trees, as well as non-exchangeable models such as the alpha-theta model, the alpha-gamma model and all restricted exchangeable models previously studied. Our main structural result is a representation of the growth rule by a sigma-finite dislocation measure kappa on the set of partitions of the natural numbers extending Bertoin's notion of exchangeable dislocation measures from the setting of homogeneous fragmentations. We use this representation to establish necessary and sufficient conditions on the growth rule under which we can apply results by Haas and Miermont for unlabelled and not necessarily consistent trees to establish self-similar random trees and residual mass processes as scaling limits. While previous studies exploited some form of exchangeability, our scaling limit results here only require a regularity condition on the convergence of asymptotic frequencies under kappa, in addition to a regular variation condition.

math.PR↗