SearcharxivSearch

arXiv subjects

Gal Yehuda

Publications and source records attributed to Gal Yehuda.

14 recordsLinked to original sources

Growth gaps and generating sets

We show that the existence of a growth gap for infinite-index subgroups of a given finitely genrated group can depend on the finite generating set. More precisely, for any irreducible lattice $\Lambda$ in a higher rank semisimple Lie group $G$ with Kazhdan's property (T), the group $\Lambda \times \Lambda$ admits one finite symmetric generating set with a growth gap and another without a growth gap. We also prove that the growth gap can be made arbitrarily small. In contrast, for a non-elementary hyperbolic group the existence of a growth gap is independent of the finite generating set.

math.GR

Improved Almost laws for $SO(3)$

We construct quantitative almost laws for $SO(3)$. More precisely, there exist a constant $c>0$ and non-trivial words $W_n\in F_2$ such that, for every $A,B\in SO(3)$, \[ \|W_n(A,B)-I\| \le \exp\!\left(-c |W_n|^{\delta}\right), \] where $\delta=\log_2(x_0)=0.879146\ldots$ and $x_0>1$ is the real root of $x^3=x^2+x+1$. This improves the exponent $\log_2\varphi$ obtained from Elkasapy's lower-central-series construction. As an application, we show how this result improves the word-length threshold in Kuperberg's Solovay--Kitaev algorithm for single-qubit gates.

math.GR

On the growth spectrum of hyperbolic groups

We study the growth spectrum of groups acting on hyperbolic spaces, i.e.\ the set of exponential growth rates achieved by subgroups. For a finitely generated free group or a surface group acting convex-cocompactly on a proper geodesic hyperbolic metric space, we prove that the growth spectrum is the full interval $[0, \omega_G]$. For any hyperbolic group, we prove that the growth spectrum contains a large interval $[0, \omega_{\mathcal{F}}]$ where $\omega_{\mathcal{F}} \geq \omega_G / 2$, with strict inequality when the action is divergent. In the case of the Cayley graph of a free group, we also present an approach via the non-backtracking matrix of the configuration model, connecting the density of growth rates to a spectral concentration result for random graphs.

math.GR

Growth of Approximate Groups in Hyperbolic Groups

We prove a growth dichotomy for infinite approximate groups, and more generally approximate semigroups, in hyperbolic groups. If \(G\) is a finitely generated hyperbolic group and \(A\subseteq G\) is infinite with \[ A^2\subseteq AX \] for some finite \(X\subseteq G\), then either \(\langle A\rangle\) is virtually cyclic, or \(A\) has positive exponential growth in the ambient word metric. We also introduce a product-growth criterion for the existence of growth rates of approximate semigroups. The criterion applies to hyperbolic groups: if \(G\) is hyperbolic with finite generating set \(S\), then there is a constant \(c_{G,S}>0\) such that \[ |UV| \geq c_{G,S}\,\frac{|U||V|}{n+k+1}, \qquad U\subseteq B_n,\; V\subseteq B_k. \] The linear loss is optimal in order whenever \(G\) contains an element of infinite order. In the free group with its standard generating set one may take \(c_{G,S}=1/4\). We also prove that, in a free group, if \(U\subseteq S_n\) and \(V\subseteq S_k\), then \[ |UV|\geq \left(\frac{2}{3}+\frac{1}{3\cdot 4^{\min\{n,k\}}}\right)|U||V|, \] and this constant is sharp for all \(n,k\).

math.GR

On the spectral radius of the non-backtracking matrix of the configuration model

We prove a concentration result for the leading eigenvalue of the non--backtracking matrix of the configuration model under the assumption of uniformly bounded degrees. Let $P$ denote the limiting degree distribution. Assuming polynomial approximation, we show that as the number of vertices tends to infinity, the leading eigenvalue of the non--backtracking matrix concentrates around \[ \frac{\mathbb{E}[P(P-1)]}{\mathbb{E}[P]}. \] This quantity corresponds to the mean offspring number of the excess--degree branching process associated with the local limit of the configuration model. As a byproduct of our work we explain how this result can be applied to prove the density of the growth rates of the subgroups of the free group.

math.GR

Geometric Covering using Random Fields

A set of vectors $S \subseteq \mathbb{R}^d$ is $(k_1,\varepsilon)$-clusterable if there are $k_1$ balls of radius $\varepsilon$ that cover $S$. A set of vectors $S \subseteq \mathbb{R}^d$ is $(k_2,δ)$-far from being clusterable if there are at least $k_2$ vectors in $S$, with all pairwise distances at least $δ$. We propose a probabilistic algorithm to distinguish between these two cases. Our algorithm reaches a decision by only looking at the extreme values of a scalar valued hash function, defined by a random field, on $S$; hence, it is especially suitable in distributed and online settings. An important feature of our method is that the algorithm is oblivious to the number of vectors: in the online setting, for example, the algorithm stores only a constant number of scalars, which is independent of the stream length. We introduce random field hash functions, which are a key ingredient in our paradigm. Random field hash functions generalize locality-sensitive hashing (LSH). In addition to the LSH requirement that ``nearby vectors are hashed to similar values", our hash function also guarantees that the ``hash values are (nearly) independent random variables for distant vectors". We formulate necessary conditions for the kernels which define the random fields applied to our problem, as well as a measure of kernel optimality, for which we provide a bound. Then, we propose a method to construct kernels which approximate the optimal one.

cs.DS

Probabilistic Invariant Learning with Randomized Linear Classifiers

Designing models that are both expressive and preserve known invariances of tasks is an increasingly hard problem. Existing solutions tradeoff invariance for computational or memory resources. In this work, we show how to leverage randomness and design models that are both expressive and invariant but use less resources. Inspired by randomized algorithms, our key insight is that accepting probabilistic notions of universal approximation and invariance can reduce our resource requirements. More specifically, we propose a class of binary classification models called Randomized Linear Classifiers (RLCs). We give parameter and sample size conditions in which RLCs can, with high probability, approximate any (smooth) function while preserving invariance to compact group transformations. Leveraging this result, we design three RLCs that are provably probabilistic invariant for classification tasks over sets, graphs, and spherical data. We show how these models can achieve probabilistic invariance and universality using less resources than (deterministic) neural networks and their invariant counterparts. Finally, we empirically demonstrate the benefits of this new class of models on invariant tasks where deterministic invariant neural networks are known to struggle.

cs.LG

Coin Flipping Neural Networks

We show that neural networks with access to randomness can outperform deterministic networks by using amplification. We call such networks Coin-Flipping Neural Networks, or CFNNs. We show that a CFNN can approximate the indicator of a $d$-dimensional ball to arbitrary accuracy with only 2 layers and $\mathcal{O}(1)$ neurons, where a 2-layer deterministic network was shown to require $Ω(e^d)$ neurons, an exponential improvement (arXiv:1610.09887). We prove a highly non-trivial result, that for almost any classification problem, there exists a trivially simple network that solves it given a sufficiently powerful generator for the network's weights. Combining these results we conjecture that for most classification problems, there is a CFNN which solves them with higher accuracy or fewer neurons than any deterministic network. Finally, we verify our proofs experimentally using novel CFNN architectures on CIFAR10 and CIFAR100, reaching an improvement of 9.25\% from the baseline.

cs.LG

A lower bound for essential covers of the cube

Essential covers were introduced by Linial and Radhakrishnan as a model that captures two complementary properties: (1) all variables must be included and (2) no element is redundant. In their seminal paper, they proved that every essential cover of the $n$-dimensional hypercube must be of size at least $Ω(n^{0.5})$. Later on, this notion found several applications in complexity theory. We improve the lower bound to $Ω(n^{0.52})$, and describe two applications.

math.CO

Slicing the hypercube is not easy

We prove that at least $Ω(n^{0.51})$ hyperplanes are needed to slice all edges of the $n$-dimensional hypercube. We provide a couple of applications: lower bounds on the computational complexity of parity, and a lower bound on the cover number of the hypercube by skew hyperplanes.

math.CO

It's Not What Machines Can Learn, It's What We Cannot Teach

Can deep neural networks learn to solve any task, and in particular problems of high complexity? This question attracts a lot of interest, with recent works tackling computationally hard tasks such as the traveling salesman problem and satisfiability. In this work we offer a different perspective on this question. Given the common assumption that $\textit{NP} \neq \textit{coNP}$ we prove that any polynomial-time sample generator for an $\textit{NP}$-hard problem samples, in fact, from an easier sub-problem. We empirically explore a case study, Conjunctive Query Containment, and show how common data generation techniques generate biased datasets that lead practitioners to over-estimate model accuracy. Our results suggest that machine learning approaches that require training on a dense uniform sampling from the target distribution cannot be used to solve computationally hard problems, the reason being the difficulty of generating sufficiently large and unbiased training sets.

cs.LG

The Complexity of Computing (Almost) Unitary Matrices With $\eps$-Copies of the Fourier Transform

The complexity of computing the Fourier transform is a longstanding open problem. Very recently, Ailon (2013, 2014, 2015) showed in a collection of papers that, roughly speaking, a speedup of the Fourier transform computation implies numerical ill-condition. The papers also quantify this tradeoff. The main method for proving these results is via a potential function called quasi-entropy, reminiscent of Shannon entropy. The quasi-entropy method opens new doors to understanding the computational complexity of the important Fourier transformation. However, it suffers from various obvious limitations. This paper, motivated by one such limitation, partly overcomes it, while at the same time sheds llight on new interesting, and problems on the intersection of computational complexity and group theory. The paper also explains why this research direction, if fruitful, has a chance of solving much bigger questions about the complexity of the Fourier transform.

cs.CC