SearcharxivSearch

arXiv subjects

Persi Diaconis

Publications and source records attributed to Persi Diaconis.

At least 19 recordsLinked to original sources

Schur--Weyl duality for diagonalizing a Markov chain on the hypercube

We show how the tools of modern algebraic combinatorics -- representation theory, Murphy elements, and particularly Schur--Weyl duality -- can be used to give an explicit orthonormal basis of eigenfunctions for a "curiously slowly mixing Markov chain" on the space of binary $n$-tuples. The basis is used to give sharp rates of convergence to stationarity.

math.RT

A curiously slowly mixing Markov chain

We study a Markov chain with very different mixing rates depending on how mixing is measured. The chain is the "Burnside process on the hypercube $C_2^n$." Started at the all-zeros state, it mixes in a bounded number of steps, no matter how large $n$ is, in $\ell^1$ and in $\ell^2$. And started at general $x$, it mixes in at most $\log n$ steps in $\ell^1$. But, in $\ell^2$, it takes $\frac{n}{\log n}$ steps for most starting $x$. The $\ell^2$ mixing results follow from an explicit diagonalization of the Markov chain into binomial-coefficient-valued eigenvectors.

math.PR

Markov chains on Weyl groups from the geometry of the flag variety

This paper studies a basic Markov chain, the Burnside process, on the space of flags $G/B$ with $G = GL_n(\mathbb{F}_q)$ and $B$ its upper triangular matrices. This gives rise to a shuffling: a Markov chain on the symmetric group realized via the Bruhat decomposition. Actually running and describing this Markov chain requires understanding Springer fibers and the Steinberg variety. The main results give a practical algorithm for all n and q and determine the limiting behavior of the chain when q is large. In describing this behavior, we find interesting connections to the combinatorics of the Robinson-Schensted correspondence and to the geometry of orbital varieties. The construction and description is then carried over to finite Chevalley groups of arbitrary type, describing a new class of Markov chains on Weyl groups.

math.PR

Permuton and local limits for the Luce model

We investigate the asymptotic properties of permutations drawn from the Luce model, a natural probabilistic framework in which permutations are generated sequentially by sampling without replacement, with selection probabilities proportional to prescribed positive weights. These permutations arise in applications such as ranking models, the Tsetlin library, and related Markov processes. Under minimal assumptions on the weights, we establish a permuton limit theorem describing the global behavior of Luce-distributed permutations and derive an explicit density of the limiting permuton. We further compute limiting pattern densities and analyze the differences between exact Luce permutations and their permuton approximations. We also study the local convergence of these permutations, proving a quenched Benjamini--Schramm limit and a central limit theorem for consecutive pattern occurrences. Finally, we prove a central limit theorem for the number of inversions.

math.PR

Estimating the size of a set using cascading exclusion

Let $S$ be a finite set, and $X_1,\ldots,X_n$ an i.i.d. uniform sample from $S$. To estimate the size $|S|$, without further structure, one can wait for repeats and use the birthday problem. This requires a sample size of the order $|S|^\frac{1}{2}$. On the other hand, if $S=\{1,2,\ldots,|S|\}$, the maximum of the sample blown up by $n/(n-1)$ gives an efficient estimator based on any growing sample size. This paper gives refinements that interpolate between these extremes. A general non-asymptotic theory is developed. This includes estimating the volume of a compact convex set, the unseen species problem, and a host of testing problems that follow from the question `Is this new observation a typical pick from a large prespecified population?' We also treat regression style predictors. A general theorem gives non-parametric finite $n$ error bounds in all cases.

math.ST

On the number and sizes of double cosets of Sylow subgroups of the symmetric group

Let $P_n$ be a Sylow $p$-subgroup of the symmetric group $S_n$. We investigate the number and sizes of the $P_n\setminus S_n\ /\ P_n$ double cosets, showing that most double cosets have maximal size when $p$ is odd, or equivalently, that $P_n\cap P_n^x=1$ for most $x\in S_n$ when $n$ is large. We also find that all possible sizes of such double cosets occur, modulo a list of small exceptions.

math.GR

Random sampling of contingency tables and partitions: Two practical examples of the Burnside process

This paper gives new, efficient algorithms for approximate uniform sampling of contingency tables and integer partitions. The algorithms use the Burnside process, a general algorithm for sampling a uniform orbit of a finite group acting on a finite set. We show that a technique called `lumping' can be used to derive efficient implementations of the Burnside process. For both contingency tables and partitions, the lumped processes have far lower per step complexity than the original Markov chains. We also define a second Markov chain for partitions called the reflected Burnside process. The reflected Burnside process maintains the computational advantages of the lumped process but empirically converges to the uniform distribution much more rapidly. By using the reflected Burnside process we can easily sample uniform partitions of size $10^{10}$.

stat.CO

An algorithm for uniform generation of unlabeled trees (P\'olya trees), with an extension of Cayley's formula

P\'olya trees are rooted, unlabeled trees on $n$ vertices. This paper gives an efficient, new way to generate P\'olya trees. This allows comparing typical unlabeled and labeled tree statistics and comparing asymptotic theorems with `reality'. Along the way, we give a product formula for the number of rooted labeled trees preserved by a given automorphism; this refines Cayley's formula.

math.CO

Poisson approximation for large permutation groups

Let $G_{k,n}$ be a group of permutations of $kn$ objects which permutes things independently in disjoint blocks of size $k$ and then permutes the blocks. We investigate the probabilistic and/or enumerative aspects of random elements of $G_{k,n}$. This includes novel limit theorems for fixed points, cycles of various lengths, number of cycles and inversions. The limits are compound Poisson distributions with interesting dependence structure.

math.PR

A Vershik-Kerov theorem for wreath products

Let $G_{n,k}$ be the group of permutations of $\{1,2,\ldots, kn\}$ that permutes the first $k$ symbols arbitrarily, then the next $k$ symbols and so on through the last $k$ symbols. Finally the $n$ blocks of size $k$ are permuted in an arbitrary way. For $\sigma$ chosen uniformly in $G_{n,k}$, let $L_{n,k}$ be the length of the longest increasing subsequence in $\sigma$. For $k,n$ growing, we determine that the limiting mean of $L_{n,k}$ is asymptotic to $4\sqrt{nk}$. This is different from parallel variations of the Vershik-Kerov theorem for colored permutations.

math.PR

Enumerative Theory for the Tsetlin Library

The Tsetlin library is a well-studied Markov chain on the symmetric group $S_n$. It has stationary distribution $π(σ)$ the Luce model, a nonuniform distribution on $S_n$, which appears in psychology, horse race betting, and tournament poker. Simple enumerative questions, such as ``what is the distribution of the top $k$ cards?'' or ``what is the distribution of the bottom $k$ cards?'' are long open. We settle these questions and draw attention to a host of parallel questions on the extension to the chambers of a hyperplane arrangement.

math.PR

On a Markov construction of couplings

For $N\in\mathbb{N}$, let $π_N$ be the law of the number of fixed points of a random permutation of $\{1, 2, ..., N\}$. Let $\mathcal{P}$ be a Poisson law of parameter 1.A classical result shows that $π_N$ converges to $\mathcal{P}$ for large $N$ and indeed in total variation $$\left\Vert π_N-\mathcal{P}\right\Vert_{\mathrm{tv}} \leq \frac{2^N}{(N+1)!}$$ This implies that $π_N$ and $\mathcal{P}$ can be coupled to at least this accuracy. This paper constructs such a coupling (a long open problem) using the machinery of intertwining of two Markov chains. This method shows promise for related problems of random matrix theory.

math.PR

Isomorphisms between random graphs

Consider two independent Erdős-Rényi $G(N,1/2)$ graphs. We show that with probability tending to $1$ as $N\to\infty$, the largest induced isomorphic subgraph has size either $\lfloor x_N-\varepsilon_N\rfloor$ or $\lfloor x_N+\varepsilon_N \rfloor$, where $x_N=4\log_2 N -2 \log_2 \log_2 N - 2\log_2(4/e)+1$ and $\varepsilon_N = (4\log_2 N)^{-1/2}$. Using similar techniques, we also show that if $Γ_1$ and $Γ_2$ are independent $G(n,1/2)$ and $G(N,1/2)$ random graphs, then $Γ_2$ contains an isomorphic copy of $Γ_1$ as an induced subgraph with high probability if $n\le \lfloor y_N - \varepsilon_N \rfloor$ and does not contain an isomorphic copy of $Γ_1$ as an induced subgraph with high probability if $n>\lfloor y_N+\varepsilon_N \rfloor$, where $y_N=2\log_2 N+1$ and $\varepsilon_N$ is as above.

math.CO

Double Coset Markov Chains

Let $G$ be a finite group. Let $H, K$ be subgroups of $G$ and $H \backslash G / K$ the double coset space. Let $Q$ be a probability on $G$ which is constant on conjugacy classes ($Q(s^{-1} t s) = Q(t)$). The random walk driven by $Q$ on $G$ projects to a Markov chain on $H \backslash G /K$. This allows analysis of the lumped chain using the representation theory of $G$. Examples include coagulation-fragmentation processes and natural Markov chains on contingency tables. Our main example projects the random transvections walk on $GL_n(q)$ onto a Markov chain on $S_n$ via the Bruhat decomposition. The chain on $S_n$ has a Mallows stationary distribution and interesting mixing time behavior. The projection illuminates the combinatorics of Gaussian elimination. Along the way, we give a representation of the sum of transvections in the Hecke algebra of double cosets. Some extensions and examples of double coset Markov chains with $G$ a compact group are discussed.

math.PR

Card Guessing with Partial Feedback

Consider the following experiment: a deck with $m$ copies of $n$ different card types is randomly shuffled, and a guesser attempts to guess the cards sequentially as they are drawn. Each time a guess is made, some amount of "feedback" is given. For example, one could tell the guesser the true identity of the card they just guessed (the complete feedback model) or they could be told nothing at all (the no feedback model). In this paper we explore a partial feedback model, where upon guessing a card, the guesser is only told whether or not their guess was correct. We show in this setting that, uniformly in $n$, at most $m+O(m^{3/4}\log m)$ cards can be guessed correctly in expectation. This resolves a question of Diaconis and Graham from 1981, where even the $m=2$ case was open.

math.PR

Sequential importance sampling for estimating expectations over the space of perfect matchings

This paper makes three contributions to estimating the number of perfect matching in bipartite graphs. First, we prove that the popular sequential importance sampling algorithm works in polynomial time for dense bipartite graphs. More carefully, our algorithm gives a $(1\pmε)$-approximation for the number of perfect matchings of a $λ$-dense bipartite graph, using $O(n^{\frac{1-2λ}λε^{-2}})$ samples. With size $n$ on each side and for $\frac{1}{2}>λ>0$, a $λ$-dense bipartite graph has all degrees greater than $(λ+\frac{1}{2})n$. Second, practical applications of the algorithm require many calls to matching algorithms. A novel preprocessing step is provided which makes significant improvements. Third, three applications are provided. The first is for counting Latin squares, the second is a practical way of computing the greedy algorithm for a card-guessing game with feedback, and the third is for stochastic block models. In all three examples, sequential importance sampling allows treating practical problems of reasonably large sizes.

math.PR

A random walk on the Rado graph

The Rado graph, also known as the random graph $G(\infty, p)$, is a classical limit object for finite graphs. We study natural ball walks as a way of understanding the geometry of this graph. For the walk started at $i$, we show that order $\log_2^*i$ steps are sufficient, and for infinitely many $i$, necessary for convergence to stationarity. The proof involves an application of Hardy's inequality for trees.

math.PR