SearcharxivSearch

arXiv subjects

Elizabeth Yang

Publications and source records attributed to Elizabeth Yang.

7 recordsLinked to original sources

Capacity Analysis of Vector Symbolic Architectures

Hyperdimensional computing (HDC) is a biologically-inspired framework which represents symbols with high-dimensional vectors, and uses vector operations to manipulate them. The ensemble of a particular vector space and a prescribed set of vector operations (including one addition-like for "bundling" and one outer-product-like for "binding") form a *vector symbolic architecture* (VSA). While VSAs have been employed in numerous applications and have been studied empirically, many theoretical questions about VSAs remain open. We analyze the *representation capacities* of four common VSAs: MAP-I, MAP-B, and two VSAs based on sparse binary vectors. "Representation capacity' here refers to bounds on the dimensions of the VSA vectors required to perform certain symbolic tasks, such as testing for set membership $i \in S$ and estimating set intersection sizes $|X \cap Y|$ for two sets of symbols $X$ and $Y$, to a given degree of accuracy. We also analyze the ability of a novel variant of a Hopfield network (a simple model of associative memory) to perform some of the same tasks that are typically asked of VSAs. In addition to providing new bounds on VSA capacities, our analyses establish and leverage connections between VSAs, "sketching" (dimensionality reduction) algorithms, and Bloom filters.

cs.LG

Local and global expansion in random geometric graphs

Consider a random geometric 2-dimensional simplicial complex $X$ sampled as follows: first, sample $n$ vectors $\boldsymbol{u_1},\ldots,\boldsymbol{u_n}$ uniformly at random on $\mathbb{S}^{d-1}$; then, for each triple $i,j,k \in [n]$, add $\{i,j,k\}$ and all of its subsets to $X$ if and only if $\langle{\boldsymbol{u_i},\boldsymbol{u_j}}\rangle \ge \tau, \langle{\boldsymbol{u_i},\boldsymbol{u_k}}\rangle \ge \tau$, and $\langle \boldsymbol{u_j}, \boldsymbol{u_k}\rangle \ge \tau$. We prove that for every $\varepsilon > 0$, there exists a choice of $d = \Theta(\log n)$ and $\tau = \tau(\varepsilon,d)$ so that with high probability, $X$ is a high-dimensional expander of average degree $n^\varepsilon$ in which each $1$-link has spectral gap bounded away from $\frac{1}{2}$. To our knowledge, this is the first demonstration of a natural distribution over $2$-dimensional expanders of arbitrarily small polynomial average degree and spectral link expansion better than $\frac{1}{2}$. All previously known constructions are algebraic. This distribution also furnishes an example of simplicial complexes for which the trickle-down theorem is nearly tight. En route, we prove general bounds on the spectral expansion of random induced subgraphs of arbitrary vertex transitive graphs, which may be of independent interest. For example, one consequence is an almost-sharp bound on the second eigenvalue of random $n$-vertex geometric graphs on $\mathbb{S}^{d-1}$, which was previously unknown for most $n,d$ pairs.

math.CO

Testing thresholds for high-dimensional sparse random geometric graphs

In the random geometric graph model $\mathsf{Geo}_d(n,p)$, we identify each of our $n$ vertices with an independently and uniformly sampled vector from the $d$-dimensional unit sphere, and we connect pairs of vertices whose vectors are ``sufficiently close'', such that the marginal probability of an edge is $p$. We investigate the problem of testing for this latent geometry, or in other words, distinguishing an Erd\H{o}s-R\'enyi graph $\mathsf{G}(n, p)$ from a random geometric graph $\mathsf{Geo}_d(n, p)$. It is not too difficult to show that if $d\to \infty$ while $n$ is held fixed, the two distributions become indistinguishable; we wish to understand how fast $d$ must grow as a function of $n$ for indistinguishability to occur. When $p = \frac{\alpha}{n}$ for constant $\alpha$, we prove that if $d \ge \mathrm{polylog} n$, the total variation distance between the two distributions is close to $0$; this improves upon the best previous bound of Brennan, Bresler, and Nagaraj (2020), which required $d \gg n^{3/2}$, and further our result is nearly tight, resolving a conjecture of Bubeck, Ding, Eldan, \& R\'{a}cz (2016) up to logarithmic factors. We also obtain improved upper bounds on the statistical indistinguishability thresholds in $d$ for the full range of $p$ satisfying $\frac{1}{n}\le p\le \frac{1}{2}$, improving upon the previous bounds by polynomial factors. Our analysis uses the Belief Propagation algorithm to characterize the distributions of (subsets of) the random vectors {\em conditioned on producing a particular graph}. In this sense, our analysis is connected to the ``cavity method'' from statistical physics. To analyze this process, we rely on novel sharp estimates for the area of the intersection of a random sphere cap with an arbitrary subset of the sphere, which we prove using optimal transport maps and entropy-transport inequalities on the unit sphere.

math.PR

Domain Sparsification of Discrete Distributions using Entropic Independence

We present a framework for speeding up the time it takes to sample from discrete distributions $\mu$ defined over subsets of size $k$ of a ground set of $n$ elements, in the regime $k\ll n$. We show that having estimates of marginals $\mathbb{P}_{S\sim \mu}[i\in S]$, the task of sampling from $\mu$ can be reduced to sampling from distributions $\nu$ supported on size $k$ subsets of a ground set of only $n^{1-\alpha}\cdot \operatorname{poly}(k)$ elements. Here, $1/\alpha\in [1, k]$ is the parameter of entropic independence for $\mu$. Further, the sparsified distributions $\nu$ are obtained by applying a sparse (mostly $0$) external field to $\mu$, an operation that often retains algorithmic tractability of sampling from $\nu$. This phenomenon, which we dub domain sparsification, allows us to pay a one-time cost of estimating the marginals of $\mu$, and in return reduce the amortized cost needed to produce many samples from the distribution $\mu$, as is often needed in upstream tasks such as counting and inference. For a wide range of distributions where $\alpha=\Omega(1)$, our result reduces the domain size, and as a corollary, the cost-per-sample, by a $\operatorname{poly}(n)$ factor. Examples include monomers in a monomer-dimer system, non-symmetric determinantal point processes, and partition-constrained Strongly Rayleigh measures. Our work significantly extends the reach of prior work of Anari and Derezi\'nski who obtained domain sparsification for distributions with a log-concave generating polynomial (corresponding to $\alpha=1$). As a corollary of our new analysis techniques, we also obtain a less stringent requirement on the accuracy of marginal estimates even for the case of log-concave polynomials; roughly speaking, we show that constant-factor approximation is enough for domain sparsification, improving over $O(1/k)$ relative error established in prior work.

cs.DS

High-Dimensional Expanders from Expanders

We present an elementary way to transform an expander graph into a simplicial complex where all high order random walks have a constant spectral gap, i.e., they converge rapidly to the stationary distribution. As an upshot, we obtain new constructions, as well as a natural probabilistic model to sample constant degree high-dimensional expanders. In particular, we show that given an expander graph $G$, adding self loops to $G$ and taking the tensor product of the modified graph with a high-dimensional expander produces a new high-dimensional expander. Our proof of rapid mixing of high order random walks is based on the decomposable Markov chains framework introduced by Jerrum et al.

cs.DM

Inclusion of Forbidden Minors in Random Representable Matroids

In 1984, Kelly and Oxley introduced the model of a random representable matroid $M[A_n]$ corresponding to a random matrix $A_n \in \mathbb{F}_q^{m(n) \times n}$, whose entries are drawn independently and uniformly from $\mathbb{F}_q$. Whereas properties such as rank, connectivity, and circuit size have been well-studied, forbidden minors have not yet been analyzed. Here, we investigate the asymptotic probability as $n \to \infty$ that a fixed $\mathbb{F}_q$-representable matroid $M$ is a minor of $M[A_n]$. (We always assume $m(n) \geq \text{rank}(M)$ for all sufficiently large $n$, otherwise $M$ can never be a minor of the corresponding $M[A_n]$.) When $M$ is free, we show that $M$ is asymptotically almost surely (a.a.s.) a minor of $M[A_n]$. When $M$ is not free, we show a phase transition: $M$ is a.a.s. a minor if $n - m(n) \to \infty$, but is a.a.s. not if $m(n) - n \to \infty$. In the more general settings of $m \leq n$ and $m > n$, we give lower and upper bounds, respectively, on both the asymptotic and non-asymptotic probability that $M$ is a minor of $M[A_n]$. The tools we develop to analyze matroid operations and minors of random matroids may be of independent interest. Our results directly imply that $M[A_n]$ is a.a.s. not contained in any proper, minor-closed class $\mathcal{M}$ of $\mathbb{F}_q$-representable matroids, provided: (i) $n - m(n) \to \infty$, and (ii) $m(n)$ is at least the minimum rank of any $\mathbb{F}_q$-representable forbidden minor of $\mathcal{M}$, for all sufficiently large $n$. As an application, this shows that graphic matroids are a vanishing subset of linear matroids, in a sense made precise in the paper. Our results provide an approach for applying the rich theory around matroid minors to the less-studied field of random matroids.

math.CO

Competition graphs induced by permutations

In prior work, Cho and Kim studied competition graphs arising from doubly partial orders. In this article, we consider a related problem where competition graphs are instead induced by permutations. We first show that this approach produces the same class of competition graphs as the doubly partial order. In addition, we observe that the $123$ and $132$ patterns in a permutation induce the edges in the associated competition graph. We classify the competition graphs arising from $132$-avoiding permutations and show that those graphs must avoid an induced path graph of length 3. Finally, we consider the weighted competition graph of permutations and give some initial enumerative and structural results in that setting.

math.CO