Searcharxiv⌕ Search

arXiv subjects

Kyle Luh

Publications and source records attributed to Kyle Luh.

28 records · Page 2Linked to original sources

Resilience of the Rank of Random Matrices

Let $M$ be an $n \times m$ matrix of independent Rademacher ($\pm 1$) random variables. It is well known that if $n \leq m$, then $M$ is of full rank with high probability. We show that this property is resilient to adversarial changes to $M$. More precisely, if $m \geq n + n^{1-\varepsilon/6}$, then even after changing the sign of $(1-\varepsilon)m/2$ entries, $M$ is still of full rank with high probability. Note that this is asymptotically best possible as one can easily make any two rows proportional with at most $m/2$ changes. Moreover, this theorem gives an asymptotic solution to a slightly weakened version of a conjecture made by Van Vu.

math.CO↗

On the counting problem in inverse Littlewood--Offord theory

Let $ε_1, \dotsc, ε_n$ be i.i.d. Rademacher random variables taking values $\pm 1$ with probability $1/2$ each. Given an integer vector $\boldsymbol{a} = (a_1, \dotsc, a_n)$, its concentration probability is the quantity $ρ(\boldsymbol{a}):=\sup_{x\in \mathbb{Z}}\Pr(ε_1 a_1+\dots+ε_n a_n = x)$. The Littlewood-Offord problem asks for bounds on $ρ(\boldsymbol{a})$ under various hypotheses on $\boldsymbol{a}$, whereas the inverse Littlewood-Offord problem, posed by Tao and Vu, asks for a characterization of all vectors $\boldsymbol{a}$ for which $ρ(\boldsymbol{a})$ is large. In this paper, we study the associated counting problem: How many integer vectors $\boldsymbol{a}$ belonging to a specified set have large $ρ(\boldsymbol{a})$? The motivation for our study is that in typical applications, the inverse Littlewood-Offord theorems are only used to obtain such counting estimates. Using a more direct approach, we obtain significantly better bounds for this problem than those obtained using the inverse Littlewood--Offord theorems of Tao and Vu and of Nguyen and Vu. Moreover, we develop a framework for deriving upper bounds on the probability of singularity of random discrete matrices that utilizes our counting result. To illustrate the methods, we present the first `exponential-type' (i.e., $\exp(-n^c)$ for some positive constant $c$) upper bounds on the singularity probability for the following two models: (i) adjacency matrices of dense signed random regular digraphs, for which the previous best known bound is $O(n^{-1/4})$ due to Cook; and (ii) dense row-regular $\{0,1\}$-matrices, for which the previous best known bound is $O_{C}(n^{-C})$ for any constant $C>0$ due to Nguyen.

math.CO↗

Eigenvector Delocalization for Non-Hermitian Random Matrices and Applications

Improving upon results of Rudelson and Vershynin, we establish delocalization bounds for eigenvectors of independent-entry random matrices. In particular, we show that with high probability every eigenvector is delocalized, meaning any subset of its coordinates carries an appropriate proportion of its mass. Our results hold for random matrices with genuinely complex as well as real entries. In both cases, our bounds match numerical simulations, up to lower order terms, indicating the optimality of our results. As an application of our methods, we also establish delocalization bounds for normal vectors to random hyperplanes. The proofs of our main results rely on a least singular value bound for genuinely complex rectangular random matrices, which generalizes a previous bound due to the first author, and may be of independent interest.

math.PR↗

Sparse Random Matrices have Simple Spectrum

Let $M_n$ be a class of symmetric sparse random matrices, with independent entries $M_{ij} = δ_{ij} ξ_{ij}$ for $i \leq j$. $δ_{ij}$ are i.i.d. Bernoulli random variables taking the value $1$ with probability $p \geq n^{-1+δ}$ for any constant $δ> 0$ and $ξ_{ij}$ are i.i.d. centered, subgaussian random variables. We show that with high probability this class of random matrices has simple spectrum (i.e. the eigenvalues appear with multiplicity one). We can slightly modify our proof to show that the adjacency matrix of a sparse Erdős-Rényi graph has simple spectrum for $n^{-1+δ} \leq p \leq 1- n^{-1+δ}$. These results are optimal in the exponent. The result for graphs has connections to the notorious graph isomorphism problem.

math.PR↗

Complex Random Matrices have no Real Eigenvalues

Let $ζ= ξ+ iξ'$ where $ξ, ξ'$ are iid copies of a mean zero, variance one, subgaussian random variable. Let $N_n$ be a $n \times n$ random matrix with entries that are iid copies of $ζ$. We prove that there exists a $c \in (0,1)$ such that the probability that $N_n$ has any real eigenvalues is less than $c^n$ where $c$ only depends on the subgaussian moment of $ξ$. The bound is optimal up to the value of the constant $c$. The principal component of the proof is an optimal tail bound on the least singular value of matrices of the form $M_n := M + N_n$ where $M$ is a deterministic complex matrix with the condition that $\|M\| \leq K n^{1/2}$ for some constant $K$ depending on the subgaussian moment of $ξ$. For this class of random variables, this result improves on the results of Pan-Zhou and Rudelson-Vershynin. In the proof of the tail bound, we develop an optimal small-ball probability bound for complex random variables that generalizes the Littlewood-Offord theory developed by Tao-Vu and Rudelson-Vershynin.

math.PR↗

Embedding large graphs into a random graph

In this paper we consider the problem of embedding almost-spanning, bounded degree graphs in a random graph. In particular, let $Δ\geq 5$, $\varepsilon > 0$ and let $H$ be a graph on $(1-\varepsilon)n$ vertices and with maximum degree $Δ$. We show that a random graph $G_{n,p}$ with high probability contains a copy of $H$, provided that $p\gg (n^{-1}\log^{1/Δ}n)^{2/(Δ+1)}$. Our assumption on $p$ is optimal up to the $polylog$ factor. We note that this $polylog$ term matches the conjectured threshold for the spanning case.

math.CO↗

Optimal Threshold for a Random Graph to be 2-Universal

For a family of graphs $\mathcal{F}$, a graph $G$ is $\mathcal{F}$-universal if $G$ contains every graph in $\mathcal{F}$ as a (not necessarily induced) subgraph. For the family of all graphs on $n$ vertices and of maximum degree at most two, $\mathcal{H}(n,2)$, we prove that there exists a constant $C$ such that for $p \geq C \left( \frac{\log n}{n^2} \right)^{\frac{1}{3}}$, the binomial random graph $G(n,p)$ is typically $\mathcal{H}(n,2)$-universal. This bound is optimal up to the constant factor as illustrated in the seminal work of Johansson, Kahn, and Vu for triangle factors. Our result improves significantly on the previous best bound of $p \geq C \left(\frac{\log n}{n}\right)^{\frac{1}{2}}$ due to Kim and Lee. In fact, we prove the stronger result that for the family of all graphs on $n$ vertices, of maximum degree at most two and of girth at least $\ell$, $\mathcal{H}^{\ell}(n,2)$, $G(n,p)$ is typically $\mathcal H^{\ell}(n,2)$-universal when $p \geq C \left(\frac{\log n}{n^{\ell -1}}\right)^{\frac{1}{\ell}}$. This result is also optimal up to the constant factor. Our results verify (in a weak form) a classical conjecture of Kahn and Kalai.

math.CO↗

Packing Loose Hamilton Cycles

A subset $C$ of edges in a $k$-uniform hypergraph $H$ is a \emph{loose Hamilton cycle} if $C$ covers all the vertices of $H$ and there exists a cyclic ordering of these vertices such that the edges in $C$ are segments of that order and such that every two consecutive edges share exactly one vertex. The binomial random $k$-uniform hypergraph $H^k_{n,p}$ has vertex set $[n]$ and an edge set $E$ obtained by adding each $k$-tuple $e\in \binom{[n]}{k}$ to $E$ with probability $p$, independently at random. Here we consider the problem of finding edge-disjoint loose Hamilton cycles covering all but $o(|E|)$ edges, referred to as the \emph{packing problem}. While it is known that the threshold probability for the appearance of a loose Hamilton cycle in $H^k_{n,p}$ is $p=Θ\left(\frac{\log n}{n^{k-1}}\right)$, the best known bounds for the packing problem are around $p=\text{polylog}(n)/n$. Here we make substantial progress and prove the following asymptotically (up to a polylog$(n)$ factor) best possible result: For $p\geq \log^{C}n/n^{k-1}$, a random $k$-uniform hypergraph $H^k_{n,p}$ with high probability contains $N:=(1-o(1))\frac{\binom{n}{k}p}{n/(k-1)}$ edge-disjoint loose Hamilton cycles. Our proof utilizes and modifies the idea of "online sprinkling" recently introduced by Vu and the first author.

math.CO↗

Dictionary Learning with Few Samples and Matrix Concentration

Let $A$ be an $n \times n$ matrix, $X$ be an $n \times p$ matrix and $Y = AX$. A challenging and important problem in data analysis, motivated by dictionary learning and other practical problems, is to recover both $A$ and $X$, given $Y$. Under normal circumstances, it is clear that this problem is underdetermined. However, in the case when $X$ is sparse and random, Spielman, Wang and Wright showed that one can recover both $A$ and $X$ efficiently from $Y$ with high probability, given that $p$ (the number of samples) is sufficiently large. Their method works for $p \ge C n^2 \log^ 2 n$ and they conjectured that $p \ge C n \log n$ suffices. The bound $n \log n$ is sharp for an obvious information theoretical reason. In this paper, we show that $p \ge C n \log^4 n$ suffices, matching the conjectural bound up to a polylogarithmic factor. The core of our proof is a theorem concerning $l_1$ concentration of random matrices, which is of independent interest. Our proof of the concentration result is based on two ideas. The first is an economical way to apply the union bound. The second is a refined version of Bernstein's concentration inequality for the sum of independent variables. Both have nothing to do with random matrices and are applicable in general settings.

math.PR↗

Community detection using spectral clustering on sparse geosocial data

In this article we identify social communities among gang members in the Hollenbeck policing district in Los Angeles, based on sparse observations of a combination of social interactions and geographic locations of the individuals. This information, coming from LAPD Field Interview cards, is used to construct a similarity graph for the individuals. We use spectral clustering to identify clusters in the graph, corresponding to communities in Hollenbeck, and compare these with the LAPD's knowledge of the individuals' gang membership. We discuss different ways of encoding the geosocial information using a graph structure and the influence on the resulting clusterings. Finally we analyze the robustness of this technique with respect to noisy and incomplete data, thereby providing suggestions about the relative importance of quantity versus quality of collected data.

stat.AP↗