SearcharxivSearch

arXiv subjects

Paxton Turner

Publications and source records attributed to Paxton Turner.

11 recordsLinked to original sources

Testing High-dimensional Multinomials with Applications to Text Analysis

Motivated by applications in text mining and discrete distribution inference, we investigate the testing for equality of probability mass functions of $K$ groups of high-dimensional multinomial distributions. A test statistic, which is shown to have an asymptotic standard normal distribution under the null, is proposed. The optimal detection boundary is established, and the proposed test is shown to achieve this optimal detection boundary across the entire parameter space of interest. The proposed method is demonstrated in simulation studies and applied to analyze two real-world datasets to examine variation among consumer reviews of Amazon movies and diversity of statistical paper abstracts.

stat.ME

Near-optimal fitting of ellipsoids to random points

Given independent standard Gaussian points $v_1, \ldots, v_n$ in dimension $d$, for what values of $(n, d)$ does there exist with high probability an origin-symmetric ellipsoid that simultaneously passes through all of the points? This basic problem of fitting an ellipsoid to random points has connections to low-rank matrix decompositions, independent component analysis, and principal component analysis. Based on strong numerical evidence, Saunderson, Parrilo, and Willsky [Proc. of Conference on Decision and Control, pp. 6031-6036, 2013] conjecture that the ellipsoid fitting problem transitions from feasible to infeasible as the number of points $n$ increases, with a sharp threshold at $n \sim d^2/4$. We resolve this conjecture up to logarithmic factors by constructing a fitting ellipsoid for some $n = Ω( \, d^2/\mathrm{polylog}(d) \,)$, improving prior work of Ghosh et al. [Proc. of Symposium on Foundations of Computer Science, pp. 954-965, 2020] that requires $n = o(d^{3/2})$. Our proof demonstrates feasibility of the least squares construction of Saunderson et al. using a convenient decomposition of a certain non-standard random matrix and a careful analysis of its Neumann expansion via the theory of graph matrices.

cs.DS

Phase transition for detecting a small community in a large network

How to detect a small community in a large network is an interesting problem, including clique detection as a special case, where a naive degree-based $χ^2$-test was shown to be powerful in the presence of an Erdős-Renyi background. Using Sinkhorn's theorem, we show that the signal captured by the $χ^2$-test may be a modeling artifact, and it may disappear once we replace the Erdős-Renyi model by a broader network model. We show that the recent SgnQ test is more appropriate for such a setting. The test is optimal in detecting communities with sizes comparable to the whole network, but has never been studied for our setting, which is substantially different and more challenging. Using a degree-corrected block model (DCBM), we establish phase transitions of this testing problem concerning the size of the small community and the edge densities in small and large communities. When the size of the small community is larger than $\sqrt{n}$, the SgnQ test is optimal for it attains the computational lower bound (CLB), the information lower bound for methods allowing polynomial computation time. When the size of the small community is smaller than $\sqrt{n}$, we establish the parameter regime where the SgnQ test has full power and make some conjectures of the CLB. We also study the classical information lower bound (LB) and show that there is always a gap between the CLB and LB in our range of interest.

math.ST

Gaussian discrepancy: a probabilistic relaxation of vector balancing

We introduce a novel relaxation of combinatorial discrepancy called Gaussian discrepancy, whereby binary signings are replaced with correlated standard Gaussian random variables. This relaxation effectively reformulates an optimization problem over the Boolean hypercube into one over the space of correlation matrices. We show that Gaussian discrepancy is a tighter relaxation than the previously studied vector and spherical discrepancy problems, and we construct a fast online algorithm that achieves a version of the Banaszczyk bound for Gaussian discrepancy. This work also raises new questions such as the Komlós conjecture for Gaussian discrepancy, which may shed light on classical discrepancy problems.

cs.DM

A Statistical Perspective on Coreset Density Estimation

Coresets have emerged as a powerful tool to summarize data by selecting a small subset of the original observations while retaining most of its information. This approach has led to significant computational speedups but the performance of statistical procedures run on coresets is largely unexplored. In this work, we develop a statistical framework to study coresets and focus on the canonical task of nonparameteric density estimation. Our contributions are twofold. First, we establish the minimax rate of estimation achievable by coreset-based estimators. Second, we show that the practical coreset kernel density estimators are near-minimax optimal over a large class of Hölder-smooth densities.

math.ST

Efficient Interpolation of Density Estimators

We study the problem of space and time efficient evaluation of a nonparametric estimator that approximates an unknown density. In the regime where consistent estimation is possible, we use a piecewise multivariate polynomial interpolation scheme to give a computationally efficient construction that converts the original estimator to a new estimator that can be queried efficiently and has low space requirements, all without adversely deteriorating the original approximation quality. Our result gives a new statistical perspective on the problem of fast evaluation of kernel density estimators in the presence of underlying smoothness. As a corollary, we give a succinct derivation of a classical result of Kolmogorov---Tikhomirov on the metric entropy of Hölder classes of smooth functions.

math.ST

Balancing Gaussian vectors in high dimension

Motivated by problems in controlled experiments, we study the discrepancy of random matrices with continuous entries where the number of columns $n$ is much larger than the number of rows $m$. Our first result shows that if $ω(1) = m = o(n)$, a matrix with i.i.d. standard Gaussian entries has discrepancy $Θ(\sqrt{n} \, 2^{-n/m})$ with high probability. This provides sharp guarantees for Gaussian discrepancy in a regime that had not been considered before in the existing literature. Our results also apply to a more general family of random matrices with continuous i.i.d entries, assuming that $m = O(n/\log{n})$. The proof is non-constructive and is an application of the second moment method. Our second result is algorithmic and applies to random matrices whose entries are i.i.d. and have a Lipschitz density. We present a randomized polynomial-time algorithm that achieves discrepancy $e^{-Ω(\log^2(n)/m)}$ with high probability, provided that $m = O(\sqrt{\log{n}})$. In the one-dimensional case, this matches the best known algorithmic guarantees due to Karmarkar--Karp. For higher dimensions $2 \leq m = O(\sqrt{\log{n}})$, this establishes the first efficient algorithm achieving discrepancy smaller than $O( \sqrt{m} )$.

cs.DM

Efficient Reconstruction of Stochastic Pedigrees

We introduce a new algorithm called {\sc Rec-Gen} for reconstructing the genealogy or \textit{pedigree} of an extant population purely from its genetic data. We justify our approach by giving a mathematical proof of the effectiveness of {\sc Rec-Gen} when applied to pedigrees from an idealized generative model that replicates some of the features of real-world pedigrees. Our algorithm is iterative and provides an accurate reconstruction of a large fraction of the pedigree while having relatively low \emph{sample complexity}, measured in terms of the length of the genetic sequences of the population. We propose our approach as a prototype for further investigation of the pedigree reconstruction problem toward the goal of applications to real-world examples. As such, our results have some conceptual bearing on the increasingly important issue of genomic privacy.

cs.DS

Conditions for Discrete Equidecomposability of Polygons

Two rational polygons $P$ and $Q$ are said to be discretely equidecomposable if there exists a piecewise affine-unimodular bijection (equivalently, a piecewise affine-linear bijection that preserves the integer lattice $\mathbb{Z} \times \mathbb{Z}$) from $P$ to $Q$. In [TW14], we developed an invariant for rational finite discrete equidecomposability known as weight. Here we extend this program with a necessary and sufficient condition for rational finite discrete equidecomposability. We close with an algorithm for detecting and constructing equidecomposability relations between rational polygons $P$ and $Q$.

math.CO

Discrete Equidecomposability and Ehrhart Theory of Polygons

Motivated by questions from Ehrhart theory, we present new results on discrete equidecomposability. Two rational polygons $P$ and $Q$ are said to be discretely equidecomposable if there exists a piecewise affine-unimodular bijection (equivalently, a piecewise affine-linear bijection that preserves the integer lattice $\mathbb{Z} \times \mathbb{Z}$) from $P$ to $Q$. In this paper, we primarily study a particular version of this notion which we call rational finite discrete equidecomposability. We construct triangles that are Ehrhart equivalent but not rationally finitely discretely equidecomposable, thus providing a partial negative answer to a question of Haase--McAllister on whether Ehrhart equivalence implies discrete equidecomposability. Surprisingly, if we delete an edge from each of these triangles, there exists an infinite rational discrete equidecomposability relation between them. Our final section addresses the topic of infinite equidecomposability with concrete examples and a potential setting for further investigation of this phenomenon.

math.CO

Aztec Castles and the dP3 Quiver

Bipartite, periodic, planar graphs known as brane tilings can be associated to a large class of quivers. This paper will explore new algebraic properties of the well-studied del Pezzo 3 quiver and geometric properties of its corresponding brane tiling. In particular, a factorization formula for the cluster variables arising from a large class of mutation sequences (called $τ-$mutation sequences) is proven; this factorization also gives a recursion on the cluster variables produced by such sequences. We can realize these sequences as walks in a triangular lattice using a correspondence between the generators of the affine symmetric group $\tilde{A_2}$ and the mutations which generate $τ-$mutation sequences. Using this bijection, we obtain explicit formulae for the cluster that corresponds to a specific alcove in the lattice. With this lattice visualization in mind, we then express each cluster variable produced in a $τ$-mutation sequence as the sum of weighted perfect matchings of a new family of subgraphs of the dP3 brane tiling, which we call Aztec castles. Our main result generalizes previous work on a certain mutation sequence on the dP3 quiver in [Zha12], and forms part of the emerging story in combinatorics and theoretical high energy physics relating cluster variables to subgraphs of the associated brane tiling.

math.CO