SearcharxivSearch

arXiv subjects

Mia Persson

Publications and source records attributed to Mia Persson.

9 recordsLinked to original sources

Fast approximate $\ell$-center clustering in high dimensional spaces

We study the design of efficient approximation algorithms for the $\ell$-center clustering and minimum-diameter $\ell$-clustering problems in high dimensional Euclidean and Hamming spaces. Our main tool is randomized dimension reduction. First, we present a general method of reducing the dependency of the running time of a hypothetical algorithm for the $\ell$-center problem in a high dimensional Euclidean space on the dimension size. Utilizing in part this method, we provide $(2+\epsilon)$- approximation algorithms for the $\ell$-center clustering and minimum-diameter $\ell$-clustering problems in Euclidean and Hamming spaces that are substantially faster than the known $2$-approximation ones when both $\ell$ and the dimension are super-logarithmic. Next, we apply the general method to the recent fast approximation algorithms with higher approximation guarantees for the $\ell$-center clustering problem in a high dimensional Euclidean space. Finally, we provide a speed-up of the known $O(1)$-approximation method for the generalization of the $\ell$-center clustering problem to include $z$ outliers (i.e., $z$ input points can be ignored while computing the maximum distance of an input point to a center) in high dimensional Euclidean and Hamming spaces.

cs.DS

Approximate all-pairs Hamming distances and 0-1 matrix multiplication

Arslan showed that computing all-pairs Hamming distances is easily reducible to arithmetic 0-1 matrix multiplication (IPL 2018). We provide a reverse, linear-time reduction of arithmetic 0-1 matrix multiplication to computing all-pairs distances in a Hamming space. On the other hand, we present a fast randomized algorithm for approximate all-pairs distances in a Hamming space. By combining it with our reduction, we obtain also a fast randomized algorithm for approximate 0-1 matrix multiplication. Next, we present an output-sensitive randomized algorithm for a minimum spanning tree of a set of points in a generalized Hamming space, the lower is the cost of the minimum spanning tree the faster is our algorithm. Finally, we provide $(2+\epsilon)$- approximation algorithms for the $\ell$-center clustering and minimum-diameter $\ell$-clustering problems in a Hamming space $\{0,1\}^d$ that are substantially faster than the known $2$-approximation ones when both $\ell$ and $d$ are super-logarithmic.

cs.DS

Multiplication of 0-1 matrices via clustering

We study applications of clustering (in particular, the $k$-center clustering problem) in the design of efficient and practical algorithms for computing an approximate and the exact arithmetic matrix product of two 0-1 rectangular matrices with clustered rows or columns, respectively. Our results in part can be regarded as an extension of the clustering-based approach to Boolean square matrix multiplication due to Arslan and Chidri (CSC 2011). First, we provide a simple and efficient deterministic algorithm for approximate matrix product of 0-1 matrices, where the additive error is proportional to the minimum maximum radius in an $\ell$-center clustering of the rows of the first matrix or an $k$-center clustering of the columns of the second matrix. Next, we use the approximation algorithm as a preprocessing after which a query asking for the exact value of an arbitrary entry in the product matrix can be answered in time proportional to the additive error. As a consequence, we obtain a simple deterministic algorithm for the exact matrix product of 0-1 matrices. We also present an improved simple deterministic algorithm for the exact product and in addition, faster analogous randomized algorithms for an approximate and the exact matrix products of 0-1 matrices based on randomized $\ell$ and $k$-center clustering.

cs.DS

$(\min,+)$ Matrix and Vector Products for Inputs Decomposable into Few Monotone Subsequences

We study the time complexity of computing the $(\min,+)$ matrix product of two $n\times n$ integer matrices in terms of $n$ and the number of monotone subsequences the rows of the first matrix and the columns of the second matrix can be decomposed into. In particular, we show that if each row of the first matrix can be decomposed into at most $m_1$ monotone subsequences and each column of the second matrix can be decomposed into at most $m_2$ monotone subsequences such that all the subsequences are non-decreasing or all of them are non-increasing then the $(\min,+)$ product of the matrices can be computed in $O(m_1m_2n^{2.569})$ time. On the other hand, we observe that if all the rows of the first matrix are non-decreasing and all columns of the second matrix are non-increasing or {\em vice versa} then this case is as hard as the general one. Similarly, we also study the time complexity of computing the $(\min,+)$ convolution of two $n$-dimensional integer vectors in terms of $n$ and the number of monotone subsequences the two vectors can be decomposed into. We show that if the first vector can be decomposed into at most $m_1$ monotone subsequences and the second vector can be decomposed into at most $m_2$ subsequences such that all the subsequences of the first vector are non-decreasing and all the subsequences of the second vector are non-increasing or {\em vice versa} then their $(\min,+)$ convolution can be computed in $\tilde{O}(m_1m_2n^{1.5})$ time. On the other, the case when both vectors are non-decreasing or both of them are non-increasing is as hard as the general case.

cs.DS

Improved Lower Bounds for Monotone q-Multilinear Boolean Circuits

A monotone Boolean circuit is composed of OR gates, AND gates and input gates corresponding to the input variables and the Boolean constants. It is $q$-multilinear if for each its output gate $o$ and for each prime implicant $s$ of the function computed at $o$, the arithmetic version of the circuit resulting from the replacement of OR and AND gates by addition and multiplication gates, respectively, computes a polynomial at $o$ which contains a monomial including the same variables as $s$ and each of the variables in $s$ has degree at most $q$ in the monomial. First, we study the complexity of computing semi-disjoint bilinear Boolean forms in terms of the size of monotone $q$-multilinear Boolean circuits. In particular, we show that any monotone $1$-multilinear Boolean circuit computing a semi-disjoint Boolean form with $p$ prime implicants includes at least $p$ AND gates. We also show that any monotone $q$-multilinear Boolean circuit computing a semi-disjoint Boolean form with $p$ prime implicants has $\Omega(\frac p {q^4})$ size. Next, we study the complexity of the monotone Boolean function $Isol_{k,n}$ that verifies if a $k$-dimensional Boolean matrix has at least one $1$ in each line (e.g., each row and column when $k=2$), in terms of monotone $q$-multilinear Boolean circuits. We show that that any $\Sigma_3$ monotone Boolean circuit for $Isol_{k,n}$ has an exponential in $n$ size or it is not $(k-1)$-multilinear.

cs.CC

An output-sensitive algorithm for all-pairs shortest paths in directed acyclic graphs

A straightforward dynamic programming method for the single-source shortest paths problem (SSSP) in an edge-weighted directed acyclic graph (DAG) processes the vertices in a topologically sorted order. First, we similarly iterate this method alternatively in a breadth-first search sorted order and the reverse order on an input directed graph with both positive and negative real edge weights, $n$ vertices and $m$ edges. For a positive integer $t,$ after $O(t)$ iterations in $O(tm)$ time, we obtain for each vertex $v$ a path distance from the source to $v$ not exceeding that yielded by the shortest path from the source to $v$ among the so called {\em$ t+$light paths}. A directed path between two vertices is $t+$light if it contains at most $t$ more edges than the minimum edge-cardinality directed path between these vertices. After $O(n)$ iterations, we obtain an $O(nm)$-time solution to SSSP in directed graphs with real edge weights matching that of Bellman and Ford. Our main result is an output-sensitive algorithm for the all-pairs shortest paths problem (APSP) in DAGs with positive and negative real edge weights. It runs in time $O(\min \{n^ω, nm+n^2\log n\}+\sum_{v\in V}\text{indeg}(v)|\text{leaf}(T_v)|),$ where $n$ is the number of vertices, $m$ is the number of edges, $ω$ is the exponent of fast matrix multiplication, $\text{indeg}(v)$ stands for the indegree of $v,$ $T_v$ is a tree of lexicographically-first shortest directed paths from all ancestors of $v$ to $v$, and $\text{leaf}(T_v)$ is the set of leaves in $T_v.$ Finally, we discuss an extension of hypothetical improved upper time-bounds for APSP in non-negatively edge-weighted DAGs to include directed graphs with a polynomial number of large directed cycles.

cs.DS

The Snow Team Problem (Clearing Directed Subgraphs by Mobile Agents)

We study several problems of clearing subgraphs by mobile agents in digraphs. The agents can move only along directed walks of a digraph and, depending on the variant, their initial positions may be pre-specified. In general, for a given subset~$\mathcal{S}$ of vertices of a digraph $D$ and a positive integer $k$, the objective is to determine whether there is a subgraph $H=(\mathcal{V}_H,\mathcal{A}_H)$ of $D$ such that (a) $\mathcal{S} \subseteq \mathcal{V}_H$, (b) $H$ is the union of $k$ directed walks in $D$, and (c) the underlying graph of $H$ includes a Steiner tree for $\mathcal{S}$ in $D$. We provide several results on the polynomial time tractability, hardness, and parameterized complexity of the problem.

cs.DM

A fast parallel algorithm for minimum-cost small integral flows

We present a new approach to the minimum-cost integral flow problem for small values of the flow. It reduces the problem to the tests of simple multi-variate polynomials over a finite field of characteristic two for non-identity with zero. In effect, we show that a minimum-cost flow of value k in a network with n vertices, a sink and a source, integral edge capacities and positive integral edge costs polynomially bounded in n can be found by a randomized PRAM, with errors of exponentially small probability in n, running in O(k\log (kn)+\log^2 (kn)) time and using 2^{k}(kn)^{O(1)} processors. Thus, in particular, for the minimum-cost flow of value O(\log n), we obtain an RNC^2 algorithm.

cs.DC