SearcharxivSearch

arXiv subjects

Henning Sulzbach

Publications and source records attributed to Henning Sulzbach.

16 recordsLinked to original sources

Dynamical Models for Random Simplicial Complexes

We study a general model of random dynamical simplicial complexes and derive a formula for the asymptotic degree distribution. This asymptotic formula encompasses results for a number of existing models, including random Apollonian networks and the weighted random recursive tree. It also confirms results on the scale-free nature of Complex Quantum Network Manifolds in dimensions $d > 2$, and special types of Network Geometry with Flavour models studied in the physics literature by Bianconi, Rahmede [$\mathit{Sci. Rep.} \; \mathbf{5},\text{ 13979 (2015) and }\mathit{Phys. Rev. E} \; \mathbf{93},\text{ 032315 (2016)}$].

math.PR

Self-similar real trees defined as fixed-points and their geometric properties

We consider fixed-point equations for probability measures charging measured compact metric spaces that naturally yield continuum random trees. On the one hand, we study the existence/uniqueness of the fixed-points and the convergence of the corresponding iterative schemes. On the other hand, we study the geometric properties of the random measured real trees that are fixed-points, in particular their fractal properties. We obtain bounds on the Minkowski and Hausdorff dimension, that are proved tight in a number of applications, including the very classical continuum random tree, but also for the dual trees of random recursive triangulations of the disk introduced by Curien and Le Gall [Ann Probab, vol. 39, 2011]. The method happens to be especially efficient to treat cases for which the mass measure on the real tree induced by natural encodings only provides weak estimates on the Hausdorff dimension.

math.PR

A limit field for orthogonal range searches in two-dimensional random point search trees

We consider the cost of general orthogonal range queries in random quadtrees. The cost of a given query is encoded into a (random) function of four variables which characterize the coordinates of two opposite corners of the query rectangle. We prove that, when suitably shifted and rescaled, the random cost function converges uniformly in probability towards a random field that is characterized as the unique solution to a distributional fixed-point equation. We also state similar results for $2$-d trees. Our results imply for instance that the worst case query satisfies the same asymptotic estimates as a typical query, and thereby resolve an old question of Chanzy, Devroye and Zamora-Cura [\emph{Acta Inf.}, 37:355--383, 2000]

math.PR

General Edgeworth expansions with applications to profiles of random trees

We prove an asymptotic Edgeworth expansion for the profiles of certain random trees including binary search trees, random recursive trees and plane-oriented random trees, as the size of the tree goes to infinity. All these models can be seen as special cases of the one-split branching random walk for which we also provide an Edgeworth expansion. These expansions lead to new results on mode, width and occupation numbers of the trees, settling several open problems raised in Devroye and Hwang [Ann. Appl. Probab. 16(2): 886--918, 2006], Fuchs, Hwang and Neininger [Algorithmica, 46 (3--4): 367--407, 2006], and Drmota and Hwang [Adv. in Appl. Probab., 37 (2): 321--341, 2005]. The aforementioned results are special cases and corollaries of a general theorem: an Edgeworth expansion for an arbitrary sequence of random or deterministic functions $\mathbb L_n:\mathbb Z\to\mathbb R$ which converges in the mod-$ϕ$-sense. Applications to Stirling numbers of the first kind will be given in a separate paper.

math.PR

Process convergence for the complexity of Radix Selection on Markov sources

A fundamental algorithm for selecting ranks from a finite subset of an ordered set is Radix Selection. This algorithm requires the data to be given as strings of symbols over an ordered alphabet, e.g., binary expansions of real numbers. Its complexity is measured by the number of symbols that have to be read. In this paper the model of independent data identically generated from a Markov chain is considered. The complexity is studied as a stochastic process indexed by the set of infinite strings over the given alphabet. The orders of mean and variance of the complexity and, after normalization, a limit theorem with a centered Gaussian process as limit are derived. This implies an analysis for two standard models for the ranks: uniformly chosen ranks, also called grand averages, and the worst case rank complexities which are of interest in computer science. For uniform data and the asymmetric Bernoulli model (i.e. memoryless sources), we also find weak convergence for the normalized process of complexities when indexed by the ranks while for more general Markov sources these processes are not tight under the standard normalizations.

math.PR

On weighted depths in random binary search trees

Following the model introduced by Aguech, Lasmar and Mahmoud [Probab. Engrg. Inform. Sci. 21 (2007) 133-141], the weighted depth of a node in a labelled rooted tree is the sum of all labels on the path connecting the node to the root. We analyze weighted depths of nodes with given labels, the last inserted node, nodes ordered as visited by the depth first search process, the weighted path length and the weighted Wiener index in a random binary search tree. We establish three regimes of nodes depending on whether the second order behaviour of their weighted depths follows from fluctuations of the keys on the path, the depth of the nodes, or both. Finally, we investigate a random distribution function on the unit interval arising as scaling limit for weighted depths of nodes with at most one child.

math.PR

The heavy path approach to Galton-Watson trees with an application to Apollonian networks

We study the heavy path decomposition of conditional Galton-Watson trees. In a standard Galton-Watson tree conditional on its size $n$, we order all children by their subtree sizes, from large (heavy) to small. A node is marked if it is among the $k$ heaviest nodes among its siblings. Unmarked nodes and their subtrees are removed, leaving only a tree of marked nodes, which we call the $k$-heavy tree. We study various properties of these trees, including their size and the maximal distance from any original node to the $k$-heavy tree. In particular, under some moment condition, the $2$-heavy tree is with high probability larger than $cn$ for some constant $c > 0$, and the maximal distance from the $k$-heavy tree is $O(n^{1/(k+1)})$ in probability. As a consequence, for uniformly random Apollonian networks of size $n$, the expected size of the longest simple path is $Ω(n)$.

math.PR

On martingale tail sums in affine two-color urn models with multiple drawings

In two recent works, Kuba and Mahmoud (arXiv:1503.090691 and arXiv:1509.09053) introduced the family of two-color affine balanced Polya urn schemes with multiple drawings. We show that, in large-index urns (urn index between $1/2$ and $1$) and triangular urns, the martingale tail sum for the number of balls of a given color admits both a Gaussian central limit theorem as well as a law of the iterated logarithm. The laws of the iterated logarithm are new even in the standard model when only one ball is drawn from the urn in each step (except for the classical Polya urn model). Finally, we prove that the martingale limits exhibit densities (bounded under suitable assumptions) and exponentially decaying tails. Applications are given in the context of node degrees in random linear recursive trees and random circuits.

math.PR

Mode and Edgeworth expansion for the Ewens distribution and the Stirling numbers

We provide asymptotic expansions for the Stirling numbers of the first kind and, more generally, the Ewens (or Karamata-Stirling) distribution. Based on these expansions, we obtain some new results on the asymptotic properties of the mode and the maximum of the Stirling numbers and the Ewens distribution. For arbitrary $θ>0$ and for all sufficiently large $n\in\mathbb N$, the unique maximum of the Ewens probability mass function $$ \mathbb L_n(k) = \frac{θ^k}{θ(θ+1)\ldots(θ+n-1)} \genfrac{[}{]}{0pt}{}{n}{k}, \quad k=1,\ldots,n, $$ is attained at $k= \left\lfloor θ\log n + \frac{θΓ'(θ)}{Γ(θ)} - \frac 12\right\rfloor$ or $k=\left\lceil θ\log n + \frac{θΓ'(θ)}{Γ(θ)} + \frac 12\right\rceil$. We prove that the mode is $$ k=\left\lfloor θ\log n - \frac{θΓ'(θ)}{Γ(θ)}\right\rfloor $$ for a set of $n$'s of asymptotic density $1$, yet this formula is not true for infinitely many $n$'s.

math.PR

On martingale tail sums for the path length in random trees

For a martingale $(X_n)$ converging almost surely to a random variable $X$, the sequence $(X_n - X)$ is called martingale tail sum. Recently, Neininger [Random Structures Algorithms, 46 (2015), 346-361] proved a central limit theorem for the martingale tail sum of R{é}gnier's martingale for the path length in random binary search trees. Gr{ü}bel and Kabluchko [to appear in Annals of Applied Probability, (2016), arXiv 1410.0469] gave an alternative proof also conjecturing a corresponding law of the iterated logarithm. We prove the central limit theorem with convergence of higher moments and the law of the iterated logarithm for a family of trees containing binary search trees, recursive trees and plane-oriented recursive trees.

math.PR

On a functional contraction method

Methods for proving functional limit laws are developed for sequences of stochastic processes which allow a recursive distributional decomposition either in time or space. Our approach is an extension of the so-called contraction method to the space $\mathcal{C}[0,1]$ of continuous functions endowed with uniform topology and the space $\mathcal {D}[0,1]$ of càdlàg functions with the Skorokhod topology. The contraction method originated from the probabilistic analysis of algorithms and random trees where characteristics satisfy natural distributional recurrences. It is based on stochastic fixed-point equations, where probability metrics can be used to obtain contraction properties and allow the application of Banach's fixed-point theorem. We develop the use of the Zolotarev metrics on the spaces $\mathcal{C}[0,1]$ and $\mathcal{D}[0,1]$ in this context. Applications are given, in particular, a short proof of Donsker's functional limit theorem is derived and recurrences arising in the probabilistic analysis of algorithms are discussed.

math.PR

The dual tree of a recursive triangulation of the disk

In the recursive lamination of the disk, one tries to add chords one after another at random; a chord is kept and inserted if it does not intersect any of the previously inserted ones. Curien and Le Gall [Ann. Probab. 39 (2011) 2224-2270] have proved that the set of chords converges to a limit triangulation of the disk encoded by a continuous process $\mathscr{M}$. Based on a new approach resembling ideas from the so-called contraction method in function spaces, we prove that, when properly rescaled, the planar dual of the discrete lamination converges almost surely in the Gromov-Hausdorff sense to a limit real tree $\mathscr{T}$, which is encoded by $\mathscr{M}$. This confirms a conjecture of Curien and Le Gall.

math.PR

Analysis of radix selection on Markov sources

The complexity of the algorithm Radix Selection is considered for independent data generated from a Markov source. The complexity is measured by the number of bucket operations required and studied as a stochastic process indexed by the ranks; also the case of a uniformly chosen rank is considered. The orders of mean and variance of the complexity and limit theorems are derived. We find weak convergence of the appropriately normalized complexity towards a Gaussian process with explicit mean and covariance functions (in the space D[0,1] of cadlag functions on [0,1] with the Skorokhod metric) for uniform data and the asymmetric Bernoulli model. For uniformly chosen ranks and uniformly distributed data the normalized complexity was known to be asymptotically normal. For a general Markov source (excluding the uniform case) we find that this complexity is less concentrated and admits a limit law with non-normal limit distribution.

math.PR

A limit process for partial match queries in random quadtrees and $2$-d trees

We consider the problem of recovering items matching a partially specified pattern in multidimensional trees (quadtrees and $k$-d trees). We assume the traditional model where the data consist of independent and uniform points in the unit square. For this model, in a structure on $n$ points, it is known that the number of nodes $C_n(ξ)$ to visit in order to report the items matching a random query $ξ$, independent and uniformly distributed on $[0,1]$, satisfies $\mathbf {E}[{C_n(ξ)}]\simκn^β$, where $κ$ and $β$ are explicit constants. We develop an approach based on the analysis of the cost $C_n(s)$ of any fixed query $s\in[0,1]$, and give precise estimates for the variance and limit distribution of the cost $C_n(x)$. Our results permit us to describe a limit process for the costs $C_n(x)$ as $x$ varies in $[0,1]$; one of the consequences is that $\mathbf {E}[{\max_{x\in[0,1]}C_n(x)}]\sim γn^β$; this settles a question of Devroye [Pers. Comm., 2000].

math.PR

A Gaussian limit process for optimal FIND algorithms

We consider versions of the FIND algorithm where the pivot element used is the median of a subset chosen uniformly at random from the data. For the median selection we assume that subsamples of size asymptotic to $c \cdot n^α$ are chosen, where $0<α\le \frac{1}{2}$, $c>0$ and $n$ is the size of the data set to be split. We consider the complexity of FIND as a process in the rank to be selected and measured by the number of key comparisons required. After normalization we show weak convergence of the complexity to a centered Gaussian process as $n\to\infty$, which depends on $α$. The proof relies on a contraction argument for probability distributions on c{à}dl{à}g functions. We also identify the covariance function of the Gaussian limit process and discuss path and tail properties.

math.PR

Partial match queries in random quadtrees

We consider the problem of recovering items matching a partially specified pattern in multidimensional trees (quad trees and k-d trees). We assume the traditional model where the data consist of independent and uniform points in the unit square. For this model, in a structure on $n$ points, it is known that the number of nodes $C_n(ξ)$ to visit in order to report the items matching an independent and uniformly on $[0,1]$ random query $ξ$ satisfies $\Ec{C_n(ξ)}\sim κn^β$, where $κ$ and $β$ are explicit constants. We develop an approach based on the analysis of the cost $C_n(x)$ of any fixed query $x\in [0,1]$, and give precise estimates for the variance and limit distribution of the cost $C_n(x)$. Our results permit to describe a limit process for the costs $C_n(x)$ as $x$ varies in $[0,1]$; one of the consequences is that $E{\max_{x\in [0,1]} C_n(x)} \sim γn^β$.

math.PR