SearcharxivSearch

arXiv subjects

Sharath Raghvendra

Publications and source records attributed to Sharath Raghvendra.

18 recordsLinked to original sources

Computing All Optimal Partial $p$-Wasserstein Matchings on the Line

For $p \ge 1$, the $p$-Wasserstein distance measures the minimum cost of transporting probability mass between distributions, where moving unit mass between two points costs the $p$th power of their distance. For discrete distributions in one dimension, full transport is especially simple: after sorting, mass is matched in order along the line. By contrast, partial and unbalanced transport on the line remains much less understood. Recently, Chapel and Tavenard [ICLR'25] showed that, for $p=1$, all optimal partial transport plans between distributions supported on $n$ points, with uniform mass at each point, can be computed in $O(n\log n)$ time by exploiting the metric structure of the cost. For $p>1$, this structure no longer applies, and existing approaches require $\Omega(n^2)$ time. Our main contribution is an FFT-based data structure for balanced-interval transport queries, which bypasses this quadratic bottleneck and yields an $O(p\,n\log^2 n)$-time algorithm for computing all optimal partial transports on the line for every finite $p\ge 1$. We also provide an open-source C++ implementation that outperforms the state-of-the-art baseline on a range of synthetic instances. Finally, we establish a conditional lower bound for $p=\infty$: any subquadratic-time algorithm for computing all optimal partial transport plan costs on the line would violate the $(\min,+)$-Convolution Hypothesis. This separates the problem from full optimal transport, which is solvable in $O(n\log n)$.

cs.CG

Geometric Bipartite Matching Based Exact Algorithms for Server Problems

For any given metric space, obtaining an offline optimal solution to the classical $k$-server problem can be reduced to solving a minimum-cost partial bipartite matching between two point sets $A$ and $B$ within that metric space. For $d$-dimensional $\ell_p$ metric space, we present an $\tilde{O}(\min\{nk, n^{2-\frac{1}{2d+1}}\log \Delta\}\cdot \Phi(n))$ time algorithm for solving this instance of minimum-cost partial bipartite matching; here, $\Delta$ represents the spread of the point set, and $\Phi(n)$ is the query/update time of a $d$-dimensional dynamic weighted nearest neighbor data structure. Our algorithm improves upon prior algorithms that require at least $\Omega(nk\Phi(n))$ time. The design of minimum-cost (partial) bipartite matching algorithms that make sub-quadratic queries to a weighted nearest-neighbor data structure, even for bounded spread instances, is a major open problem in computational geometry. We resolve this problem at least for the instances that are generated by the offline version of the $k$-server problem. Our algorithm employs a hierarchical partitioning approach, dividing the points of $A\cup B$ into rectangles. It maintains a minimum-cost partial matching where any point $b \in B$ is either matched to a point $a\in A$ or to the boundary of the rectangle it is located in. The algorithm involves iteratively merging pairs of rectangles by erasing the shared boundary between them and recomputing the minimum-cost partial matching. This continues until all boundaries are erased and we obtain the desired minimum-cost partial matching of $A$ and $B$. We exploit geometry in our analysis to show that each point participates in only $\tilde{O}(n^{1-\frac{1}{2d+1}}\log \Delta)$ number of augmenting paths, leading to a total execution time of $\tilde{O}(n^{2-\frac{1}{2d+1}}\Phi(n)\log \Delta)$.

cs.CG

A New Robust Partial $p$-Wasserstein-Based Metric for Comparing Distributions

The $2$-Wasserstein distance is sensitive to minor geometric differences between distributions, making it a very powerful dissimilarity metric. However, due to this sensitivity, a small outlier mass can also cause a significant increase in the $2$-Wasserstein distance between two similar distributions. Similarly, sampling discrepancy can cause the empirical $2$-Wasserstein distance on $n$ samples in $\mathbb{R}^2$ to converge to the true distance at a rate of $n^{-1/4}$, which is significantly slower than the rate of $n^{-1/2}$ for $1$-Wasserstein distance. We introduce a new family of distances parameterized by $k \ge 0$, called $k$-RPW that is based on computing the partial $2$-Wasserstein distance. We show that (1) $k$-RPW satisfies the metric properties, (2) $k$-RPW is robust to small outlier mass while retaining the sensitivity of $2$-Wasserstein distance to minor geometric differences, and (3) when $k$ is a constant, $k$-RPW distance between empirical distributions on $n$ samples in $\mathbb{R}^2$ converges to the true distance at a rate of $n^{-1/3}$, which is faster than the convergence rate of $n^{-1/4}$ for the $2$-Wasserstein distance. Using the partial $p$-Wasserstein distance, we extend our distance to any $p \in [1,\infty]$. By setting parameters $k$ or $p$ appropriately, we can reduce our distance to the total variation, $p$-Wasserstein, and the Lévy-Prokhorov distances. Experiments show that our distance function achieves higher accuracy in comparison to the $1$-Wasserstein, $2$-Wasserstein, and TV distances for image retrieval tasks on noisy real-world data sets.

cs.LG

Fast and Accurate Approximations of the Optimal Transport in Semi-Discrete and Discrete Settings

Given a $d$-dimensional continuous (resp. discrete) probability distribution $μ$ and a discrete distribution $ν$, the semi-discrete (resp. discrete) Optimal Transport (OT) problem asks for computing a minimum-cost plan to transport mass from $μ$ to $ν$; we assume $n$ to be the size of the support of the discrete distributions, and we assume we have access to an oracle outputting the mass of $μ$ inside a constant-complexity region in $O(1)$ time. In this paper, we present three approximation algorithms for the OT problem. (i) Semi-discrete additive approximation: For any $ε>0$, we present an algorithm that computes a semi-discrete transport plan with $ε$-additive error in $n^{O(d)}\log\frac{C_{\max}}ε$ time; here, $C_{\max}$ is the diameter of the supports of $μ$ and $ν$. (ii) Semi-discrete relative approximation: For any $ε>0$, we present an algorithm that computes a $(1+ε)$-approximate semi-discrete transport plan in $nε^{-O(d)}\log(n)\log^{O(d)}(\log n)$ time; here, we assume the ground distance is any $L_p$ norm. (iii) Discrete relative approximation: For any $ε>0$, we present a Monte-Carlo $(1+ε)$-approximation algorithm that computes a transport plan under any $L_p$ norm in $nε^{-O(d)}\log(n)\log^{O(d)}(\log n)$ time; here, we assume that the spread of the supports of $μ$ and $ν$ is polynomially bounded.

cs.CG

Deterministic, Near-Linear $\varepsilon$-Approximation Algorithm for Geometric Bipartite Matching

Given point sets $A$ and $B$ in $\mathbb{R}^d$ where $A$ and $B$ have equal size $n$ for some constant dimension $d$ and a parameter $\varepsilon>0$, we present the first deterministic algorithm that computes, in $n\cdot(\varepsilon^{-1} \log n)^{O(d)}$ time, a perfect matching between $A$ and $B$ whose cost is within a $(1+\varepsilon)$ factor of the optimal under any $\smash{\ell_p}$-norm. Although a Monte-Carlo algorithm with a similar running time is proposed by Raghvendra and Agarwal [J. ACM 2020], the best-known deterministic $\varepsilon$-approximation algorithm takes $Ω(n^{3/2})$ time. Our algorithm constructs a (refinement of a) tree cover of $\mathbb{R}^d$, and we develop several new tools to apply a tree-cover based approach to compute an $\varepsilon$-approximate perfect matching.

cs.DS

A Push-Relabel Based Additive Approximation for Optimal Transport

Optimal Transport is a popular distance metric for measuring similarity between distributions. Exact algorithms for computing Optimal Transport can be slow, which has motivated the development of approximate numerical solvers (e.g. Sinkhorn method). We introduce a new and very simple combinatorial approach to find an $\varepsilon$-approximation of the OT distance. Our algorithm achieves a near-optimal execution time of $O(n^2/\varepsilon^2)$ for computing OT distance and, for the special case of the assignment problem, the execution time improves to $O(n^2/\varepsilon)$. Our algorithm is based on the push-relabel framework for min-cost flow problems. Unlike the other combinatorial approach (Lahn, Mulchandani and Raghvendra, NeurIPS 2019) which does not have a fast parallel implementation, our algorithm has a parallel execution time of $O(\log n/\varepsilon^2)$. Interestingly, unlike the Sinkhorn algorithm, our method also readily provides a compact transport plan as well as a solution to an approximate version of the dual formulation of the OT problem, both of which have numerous applications in Machine Learning. For the assignment problem, we provide both a CPU implementation as well as an implementation that exploits GPU parallelism. Experiments suggest that our algorithm is faster than the Sinkhorn algorithm, both in terms of CPU and GPU implementations, especially while computing matchings with a high accuracy.

cs.LG

Improved Approximate Rips Filtrations with Shifted Integer Lattices and Cubical Complexes

Rips complexes are important structures for analyzing topological features of metric spaces. Unfortunately, generating these complexes is expensive because of a combinatorial explosion in the complex size. For $n$ points in $\mathbb{R}^d$, we present a scheme to construct a $2$-approximation of the filtration of the Rips complex in the $L_\infty$-norm, which extends to a $2d^{0.25}$-approximation in the Euclidean case. The $k$-skeleton of the resulting approximation has a total size of $n2^{O(d\log k +d)}$. The scheme is based on the integer lattice and simplicial complexes based on the barycentric subdivision of the $d$-cube. We extend our result to use cubical complexes in place of simplicial complexes by introducing cubical maps between complexes. We get the same approximation guarantee as the simplicial case, while reducing the total size of the approximation to only $n2^{O(d)}$ (cubical) cells. There are two novel techniques that we use in this paper. The first is the use of acyclic carriers for proving our approximation result. In our application, these are maps which relate the Rips complex and the approximation in a relatively simple manner and greatly reduce the complexity of showing the approximation guarantee. The second technique is what we refer to as scale balancing, which is a simple trick to improve the approximation ratio under certain conditions.

cs.CG

An $\tilde{O}(n^{5/4})$ Time $\varepsilon$-Approximation Algorithm for RMS Matching in a Plane

The 2-Wasserstein distance (or RMS distance) is a useful measure of similarity between probability distributions that has exciting applications in machine learning. For discrete distributions, the problem of computing this distance can be expressed in terms of finding a minimum-cost perfect matching on a complete bipartite graph given by two multisets of points $A,B \subset \mathbb{R}^2$, with $|A|=|B|=n$, where the ground distance between any two points is the squared Euclidean distance between them. Although there is a near-linear time relative $\varepsilon$-approximation algorithm for the case where the ground distance is Euclidean (Sharathkumar and Agarwal, JACM 2020), all existing relative $\varepsilon$-approximation algorithms for the RMS distance take $Ω(n^{3/2})$ time. This is primarily because, unlike Euclidean distance, squared Euclidean distance is not a metric. In this paper, for the RMS distance, we present a new $\varepsilon$-approximation algorithm that runs in $O(n^{5/4}\mathrm{poly}\{\log n,1/\varepsilon\})$ time. Our algorithm is inspired by a recent approach for finding a minimum-cost perfect matching in bipartite planar graphs (Asathulla et al., TALG 2020). Their algorithm depends heavily on the existence of sub-linear sized vertex separators as well as shortest path data structures that require planarity. Surprisingly, we are able to design a similar algorithm for a complete geometric graph that is far from planar and does not have any vertex separators. Central components of our algorithm include a quadtree-based distance that approximates the squared Euclidean distance and a data structure that supports both Hungarian search and augmentation in sub-linear time.

cs.CG

A Graph Theoretic Additive Approximation of Optimal Transport

Transportation cost is an attractive similarity measure between probability distributions due to its many useful theoretical properties. However, solving optimal transport exactly can be prohibitively expensive. Therefore, there has been significant effort towards the design of scalable approximation algorithms. Previous combinatorial results [Sharathkumar, Agarwal STOC '12, Agarwal, Sharathkumar STOC '14] have focused primarily on the design of near-linear time multiplicative approximation algorithms. There has also been an effort to design approximate solutions with additive errors [Cuturi NIPS '13, Altschuler \etal\ NIPS '17, Dvurechensky \etal\, ICML '18, Quanrud, SOSA '19] within a time bound that is linear in the size of the cost matrix and polynomial in $C/δ$; here $C$ is the largest value in the cost matrix and $δ$ is the additive error. We present an adaptation of the classical graph algorithm of Gabow and Tarjan and provide a novel analysis of this algorithm that bounds its execution time by $O(\frac{n^2 C}δ+ \frac{nC^2}{δ^2})$. Our algorithm is extremely simple and executes, for an arbitrarily small constant $\varepsilon$, only $\lfloor \frac{2C}{(1-\varepsilon)δ}\rfloor + 1$ iterations, where each iteration consists only of a Dijkstra-type search followed by a depth-first search. We also provide empirical results that suggest our algorithm is competitive with respect to a sequential implementation of the Sinkhorn algorithm in execution time. Moreover, our algorithm quickly computes a solution for very small values of $δ$ whereas Sinkhorn algorithm slows down due to numerical instability.

cs.LG

A Weighted Approach to the Maximum Cardinality Bipartite Matching Problem with Applications in Geometric Settings

We present a weighted approach to compute a maximum cardinality matching in an arbitrary bipartite graph. Our main result is a new algorithm that takes as input a weighted bipartite graph $G(A\cup B,E)$ with edge weights of $0$ or $1$. Let $w \leq n$ be an upper bound on the weight of any matching in $G$. Consider the subgraph induced by all the edges of $G$ with a weight $0$. Suppose every connected component in this subgraph has $\mathcal{O}(r)$ vertices and $\mathcal{O}(mr/n)$ edges. We present an algorithm to compute a maximum cardinality matching in $G$ in $\tilde{\mathcal{O}}( m(\sqrt{w}+ \sqrt{r}+\frac{wr}{n}))$ time. When all the edge weights are $1$ (symmetrically when all weights are $0$), our algorithm will be identical to the well-known Hopcroft-Karp (HK) algorithm, which runs in $\mathcal{O}(m\sqrt{n})$ time. However, if we can carefully assign weights of $0$ and $1$ on its edges such that both $w$ and $r$ are sub-linear in $n$ and $wr=\mathcal{O}(n^γ)$ for $γ< 3/2$, then we can compute maximum cardinality matching in $G$ in $o(m\sqrt{n})$ time. Using our algorithm, we obtain a new $\tilde{\mathcal{O}}(n^{4/3}/\varepsilon^4)$ time algorithm to compute an $\varepsilon$-approximate bottleneck matching of $A,B\subset\mathbb{R}^2$ and an $\frac{1}{\varepsilon^{\mathcal{O}(d)}}n^{1+\frac{d-1}{2d-1}}\mathrm{poly}\log n$ time algorithm for computing $\varepsilon$-approximate bottleneck matching in $d$-dimensions. All previous algorithms take $Ω(n^{3/2})$ time. Given any graph $G(A \cup B,E)$ that has an easily computable balanced vertex separator for every subgraph $G'(V',E')$ of size $|V'|^δ$, for $δ\in [1/2,1)$, we can apply our algorithm to compute a maximum matching in $\tilde{\mathcal{O}}(mn^{\fracδ{1+δ}})$ time improving upon the $\mathcal{O}(m\sqrt{n})$ time taken by the HK-Algorithm.

cs.CG

Improved Topological Approximations by Digitization

Čech complexes are useful simplicial complexes for computing and analyzing topological features of data that lies in Euclidean space. Unfortunately, computing these complexes becomes prohibitively expensive for large-sized data sets even for medium-to-low dimensional data. We present an approximation scheme for $(1+ε)$-approximating the topological information of the Čech complexes for $n$ points in $\mathbb{R}^d$, for $ε\in(0,1]$. Our approximation has a total size of $n\left(\frac{1}ε\right)^{O(d)}$ for constant dimension $d$, improving all the currently available $(1+ε)$-approximation schemes of simplicial filtrations in Euclidean space. Perhaps counter-intuitively, we arrive at our result by adding additional $n\left(\frac{1}ε\right)^{O(d)}$ sample points to the input. We achieve a bound that is independent of the spread of the point set by pre-identifying the scales at which the Čech complexes changes and sampling accordingly.

cs.CG

A Faster Algorithm for Minimum-Cost Bipartite Matching in Minor-Free Graphs

We give an $\tilde{O}(n^{7/5} \log (nC))$-time algorithm to compute a minimum-cost maximum cardinality matching (optimal matching) in $K_h$-minor free graphs with $h=O(1)$ and integer edge weights having magnitude at most $C$. This improves upon the $\tilde{O}(n^{10/7}\log{C})$ algorithm of Cohen et al. [SODA 2017] and the $O(n^{3/2}\log (nC))$ algorithm of Gabow and Tarjan [SIAM J. Comput. 1989]. For a graph with $m$ edges and $n$ vertices, the well-known Hungarian Algorithm computes a shortest augmenting path in each phase in $O(m)$ time, yielding an optimal matching in $O(mn)$ time. The Hopcroft-Karp [SIAM J. Comput. 1973], and Gabow-Tarjan [SIAM J. Comput. 1989] algorithms compute, in each phase, a maximal set of vertex-disjoint shortest augmenting paths (for appropriately defined costs) in $O(m)$ time. This reduces the number of phases from $n$ to $O(\sqrt{n})$ and the total execution time to $O(m\sqrt{n})$. In order to obtain our speed-up, we relax the conditions on the augmenting paths and iteratively compute, in each phase, a set of carefully selected augmenting paths that are not restricted to be shortest or vertex-disjoint. As a result, our algorithm computes substantially more augmenting paths in each phase, reducing the number of phases from $O(\sqrt{n})$ to $O(n^{2/5})$. By using small vertex separators, the execution of each phase takes $\tilde{O}(m)$ time on average. For planar graphs, we combine our algorithm with efficient shortest path data structures to obtain a minimum-cost perfect matching in $\tilde{O}(n^{6/5} \log{(nC)})$ time. This improves upon the recent $\tilde{O}(n^{4/3}\log{(nC)})$ time algorithm by Asathulla et al. [SODA 2018].

cs.DS

Optimal Analysis of an Online Algorithm for the Bipartite Matching Problem on a Line

In the online metric bipartite matching problem, we are given a set $S$ of server locations in a metric space. Requests arrive one at a time, and on its arrival, we need to immediately and irrevocably match it to a server at a cost which is equal to the distance between these locations. A $α$-competitive algorithm will assign requests to servers so that the total cost is at most $α$ times the cost of $M_{OPT}$ where $M_{OPT}$ is the minimum cost matching between $S$ and $R$. We consider this problem in the adversarial model for the case where $S$ and $R$ are points on a line and $|S|=|R|=n$. We improve the analysis of the deterministic Robust Matching Algorithm (RM-Algorithm, Nayyar and Raghvendra FOCS'17) from $O(\log^2 n)$ to an optimal $Θ(\log n)$. Previously, only a randomized algorithm under a weaker oblivious adversary achieved a competitive ratio of $O(\log n)$ (Gupta and Lewi, ICALP'12). The well-known Work Function Algorithm (WFA) has a competitive ratio of $O(n)$ and $Ω(\log n)$ for this problem. Therefore, WFA cannot achieve an asymptotically better competitive ratio than the RM-Algorithm.

cs.CG

Improved Approximate Rips Filtrations with Shifted Integer Lattices

Rips complexes are important structures for analyzing topological features of metric spaces. Unfortunately, generating these complexes constitutes an expensive task because of a combinatorial explosion in the complex size. For $n$ points in $\mathbb{R}^d$, we present a scheme to construct a $3\sqrt{2}$-approximation of the multi-scale filtration of the $L_\infty$-Rips complex, which extends to a $O(d^{0.25})$-approximation of the Rips filtration for the Euclidean case. The $k$-skeleton of the resulting approximation has a total size of $n2^{O(d\log k)}$. The scheme is based on the integer lattice and on the barycentric subdivision of the $d$-cube.

cs.CG

A Grid-Based Approximation Algorithm for the Minimum Weight Triangulation Problem

Given a set of $n$ points on a plane, in the Minimum Weight Triangulation problem, we wish to find a triangulation that minimizes the sum of Euclidean length of its edges. This incredibly challenging problem has been studied for more than four decades and has been only recently shown to be NP-Hard. In this paper we present a novel polynomial-time algorithm that computes a $14$-approximation of the minimum weight triangulation -- a constant that is significantly smaller than what has been previously known. In our algorithm, we use grids to partition the edges into levels where shorter edges appear at smaller levels and edges with similar lengths appear at the same level. We then triangulate the point set incrementally by introducing edges in increasing order of their levels. We introduce the edges of any level $i+1$ in two steps. In the first step, we add edges using a variant of the well-known ring heuristic to generate a partial triangulation $\hat{\mathcal{A}}_i$. In the second step, we greedily add non-intersecting level $i+1$ edges to $\hat{\mathcal{A}}_i$ in increasing order of their length and obtain a partial triangulation $\mathcal{A}_{ i+1}$. The ring heuristic is known to yield only an $O(\log n)$-approximation even for a convex polygon and the greedy heuristic achieves only a $Θ(\sqrt{n})$-approximation. Therefore, it is surprising that their combination leads to an improved approximation ratio of $14$. For the proof, we identify several useful properties of $\hat{\mathcal{A}}_i$ and combine it with a new Euler characteristic based technique to show that $\hat{\mathcal{A}}_i$ has more edges than $\mathcal{T}_i$; here $\mathcal{T}_i$ is the partial triangulation consisting of level $\le i$ edges of some minimum weight triangulation. We then use a simple greedy stays ahead proof strategy to bound the approximation ratio.

cs.CG

Polynomial-Sized Topological Approximations Using The Permutahedron

Classical methods to model topological properties of point clouds, such as the Vietoris-Rips complex, suffer from the combinatorial explosion of complex sizes. We propose a novel technique to approximate a multi-scale filtration of the Rips complex with improved bounds for size: precisely, for $n$ points in $\mathbb{R}^d$, we obtain a $O(d)$-approximation with at most $n2^{O(d \log k)}$ simplices of dimension $k$ or lower. In conjunction with dimension reduction techniques, our approach yields a $O(\mathrm{polylog} (n))$-approximation of size $n^{O(1)}$ for Rips filtrations on arbitrary metric spaces. This result stems from high-dimensional lattice geometry and exploits properties of the permutahedral lattice, a well-studied structure in discrete geometry. Building on the same geometric concept, we also present a lower bound result on the size of an approximate filtration: we construct a point set for which every $(1+ε)$-approximation of the Čech filtration has to contain $n^{Ω(\log\log n)}$ features, provided that $ε<\frac{1}{\log^{1+c} n}$ for $c\in(0,1)$.

cs.CG

Approximation and Streaming Algorithms for Projective Clustering via Random Projections

Let $P$ be a set of $n$ points in $\mathbb{R}^d$. In the projective clustering problem, given $k, q$ and norm $ρ\in [1,\infty]$, we have to compute a set $\mathcal{F}$ of $k$ $q$-dimensional flats such that $(\sum_{p\in P}d(p, \mathcal{F})^ρ)^{1/ρ}$ is minimized; here $d(p, \mathcal{F})$ represents the (Euclidean) distance of $p$ to the closest flat in $\mathcal{F}$. We let $f_k^q(P,ρ)$ denote the minimal value and interpret $f_k^q(P,\infty)$ to be $\max_{r\in P}d(r, \mathcal{F})$. When $ρ=1,2$ and $\infty$ and $q=0$, the problem corresponds to the $k$-median, $k$-mean and the $k$-center clustering problems respectively. For every $0 < ε< 1$, $S\subset P$ and $ρ\ge 1$, we show that the orthogonal projection of $P$ onto a randomly chosen flat of dimension $O(((q+1)^2\log(1/ε)/ε^3) \log n)$ will $ε$-approximate $f_1^q(S,ρ)$. This result combines the concepts of geometric coresets and subspace embeddings based on the Johnson-Lindenstrauss Lemma. As a consequence, an orthogonal projection of $P$ to an $O(((q+1)^2 \log ((q+1)/ε)/ε^3) \log n)$ dimensional randomly chosen subspace $ε$-approximates projective clusterings for every $k$ and $ρ$ simultaneously. Note that the dimension of this subspace is independent of the number of clusters~$k$. Using this dimension reduction result, we obtain new approximation and streaming algorithms for projective clustering problems. For example, given a stream of $n$ points, we show how to compute an $ε$-approximate projective clustering for every $k$ and $ρ$ simultaneously using only $O((n+d)((q+1)^2\log ((q+1)/ε))/ε^3 \log n)$ space. Compared to standard streaming algorithms with $Ω(kd)$ space requirement, our approach is a significant improvement when the number of input points and their dimensions are of the same order of magnitude.

cs.CG

Accurate Streaming Support Vector Machines

A widely-used tool for binary classification is the Support Vector Machine (SVM), a supervised learning technique that finds the "maximum margin" linear separator between the two classes. While SVMs have been well studied in the batch (offline) setting, there is considerably less work on the streaming (online) setting, which requires only a single pass over the data using sub-linear space. Existing streaming algorithms are not yet competitive with the batch implementation. In this paper, we use the formulation of the SVM as a minimum enclosing ball (MEB) problem to provide a streaming SVM algorithm based off of the blurred ball cover originally proposed by Agarwal and Sharathkumar. Our implementation consistently outperforms existing streaming SVM approaches and provides higher accuracies than libSVM on several datasets, thus making it competitive with the standard SVM batch implementation.

cs.LG