Searcharxiv⌕ Search

arXiv · 2609.38834

Optimal VC Dimension of Contrastive Learning with Margin

Abstract

Contrastive learning is a successful paradigm for learning $d$-dimensional geometric representations from a collection of ``anchor--positive--negative'' triplets $(i,j^{+},k^{-})$, indicating that ``item $i$ is closer to $j$ than to $k$.'' Despite its success, understanding why contrastive learning leads to representations of high \textit{generalization} quality---beyond the often pessimistic predictions from PAC-learning---remains a central question. Recently, \citet*{alon2024optimal} proved that, for PAC-learning $d$-dimensional Euclidean representations of $n$-point datasets, $Θ(\min(nd, n^2))$ triplets are necessary and sufficient, while they posed as an open question whether their VC dimension bounds for the more realistic setting of \textit{contrastive learning with a margin} can be improved. For a margin parameter $α>0$, a triplet $(i,j^{+},k^{-})_α$ is satisfied by the embedding $ϕ:[n]\rightarrow \mathbb{R}^{d}$, if $\|ϕ(i)-ϕ(k)\|_2>(1+α)\cdot\|ϕ(i)-ϕ(j)\|_2$. In this work, we resolve their question by proving that the VC dimension of contrastive learning under any margin $α\in(0,1)$ is in fact $O(n/α^2)$, improving on the previous bound of $O(n\log(n)/α^2)$. We also establish that the bounds are optimal up to constant factors, by providing a matching lower bound of $Ω(\frac{n}{α^2})$ (the previously known lower bound was $Ω(\frac{n}α)$), for $α\geq \max(n^{-1/2},d^{-1/2})$.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Dionysis Arvanitakis, Vaggos Chatziafratis, Yiyuan Luo, Konstantin Makarychev. 2026-09-30. Optimal VC Dimension of Contrastive Learning with Margin. https://arxiv.org/abs/2609.38834

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Auction-Based Algorithms for Matroid Intersection: Near-Linear Query Complexity and Constant-Pass Semi-Streaming

In this paper, we develop a new auction-based framework for matroid intersection and use it to obtain improved approximation algorithms in several computational settings. Our framework is inspired by Fleiner's generalized stable matching algorithm and extends the semi-streaming auction algorithm for bipartite matching due to Assadi, Liu, and Tarjan. Using this framework, for any $\varepsilon > 0$, we present a simple $(1-\varepsilon)$-approximation algorithm in the rank-oracle model whose query complexity matches that of the current fastest algorithm. Furthermore, by extending this result, we obtain the first $(1-\varepsilon)$-approximation algorithm for the weighted problem that requires only a near-linear number of rank-oracle queries, achieving the best known rank-oracle query complexity for the problem. We also obtain a $(1-\varepsilon)$-approximation semi-streaming algorithm for matroid intersection in the multi-pass streaming model, where the elements of the ground set arrive sequentially. It is the first algorithm achieving this approximation ratio using a constant number of passes and nearly linear space in the ranks of the matroids. When viewed in the standard offline setting, the same algorithm yields the first deterministic $(1-\varepsilon)$-approximation algorithm for matroid intersection that requires only a near-linear number of independence-oracle queries.

cs.DS↗

Faster network motif discovery by counting isomorphic subtrees

We develop a new algorithm for counting the number of subgraphs of a network isomorphic to a given query graph (#SubgraphIsomorphism), motivated by network motif search. High-degree vertices (hubs), common in real-world networks, contribute to a combinatorial explosion in the number of subgraphs, making existing motif search algorithms intractable for motif sizes greater than $\approx 8$ on a wide variety of networks of interest. Our procedure leverages the $k$-core decomposition and a novel subtree-counting technique to quickly scan the periphery of a network. These two innovations allow our algorithm to significantly speed up its predecessors in practice, especially as most real-world networks have a relatively large periphery. We prove that #RootedSubtreeIsomorphism, a key subroutine in our algorithm, is #P-complete via a reduction from counting bipartite matchings. We provide analytic upper bounds on our algorithm's execution time, and evaluate its performance on 11 real-world networks of varying topologies.

cs.DS↗

A Robustified Greedy Algorithm for Online Transportation with Improved Competitive Guarantees

We study the \emph{online transportation problem}, in which $n$ requests arriving sequentially in a metric space must be irrevocably assigned to $k$ capacitated facilities. Beyond classical logistics applications, this problem models resource-allocation tasks arising in machine learning, including online facility assignments, recommender systems, and mixture-of-experts routing. We introduce \emph{Robustified Greedy} (RG), a deterministic generalization of the Robust Matching algorithm that achieves a competitive ratio of $6.6604k-2.89$, improving upon the state-of-the-art bounds of $8k-7$ (Arndt et al., SOSA 2026) and $8k-5$ (Harada and Itoh, ICALP 2025). RG also retains the metric-sensitive guarantee established for Robust Matching (RM) (Nayyar and Raghvendra, FOCS 2017), achieving a competitive ratio of $O(k^{1-1/d}\log^2 n)$ in $d$-dimensional Euclidean spaces for fixed $d>1$. No comparable metric-sensitive guarantee is known for the transportation algorithms of Arndt et al.\ or Harada and Itoh. Beyond these competitive guarantees, RG provides a simple explanation for its decisions. It favors the natural nearest-neighbor assignment and, for suitable parameters, departs from this choice only when it identifies a reassignment that reduces the cost of its maintained auxiliary matching, thereby correcting accumulated assignment costs. We also prove that nearest-neighbor assignments account for a guaranteed fraction of RG's total cost, approaching one-half for appropriate parameters, even under adversarial arrivals. Experiments on real-world datasets corroborate the theory: RG achieves lower cost-to-\textsc{Opt} ratios than the competing algorithms while retaining a substantial nearest-neighbor component in its cost.

cs.DS↗