SearcharxivSearch

arXiv subjects

Noga Alon

Publications and source records attributed to Noga Alon.

At least 19 recordsLinked to original sources

Sorting from Counterexamples

Consider the following problem of learning an unknown linear order on $n$ items. In each round, the learner guesses a complete ordering of the items and receives either confirmation that the guess is correct or a counterexample: a pair of items in the wrong order. The goal is to identify the unknown order using as few queries as possible. We study this problem when up to $k$ of the returned counterexamples may be untruthful, where $k$ is not known in advance. We determine the optimal query complexity up to constant factors: \[ \Theta(n\log n + nk). \] Thus, while the noiseless complexity matches the classical complexity of sorting, each untruthful counterexample incurs an additional cost of order $n$. The upper bound is based on a geometric representation of permutations and Gr\"unbaum's theorem, while the lower bound combines sorting arguments with a Condorcet-type construction. We also study the case where the target ranking has a low-dimensional geometric representation: each item is represented by a point in $\mathbb{R}^d$, and the ranking is obtained by projecting the points onto an unknown direction. For these classes we give an upper bound of $O(d^2\log n+dk)$ and a lower bound of $\Omega(d\log n+dk)$, leaving a factor of $d$ gap in the noiseless term.

cs.LG

Pseudoshattering Pairs

For two vectors $x,y\in [b]^k$, consider the bipartite graph with two copies of $[b]$ in which $i$ on the left is joined to $j$ on the right if $(x_t,y_t)=(i,j)$ for some coordinate $t$. We study the largest size of a family $C\subseteq [b]^k$ such that, for every two distinct $x,y\in C$, this bipartite graph contains a cycle. We give a natural construction for such families and conjecture that it is optimal whenever $k$ is large relative to $b$. We prove an LYM-type upper bound that is asymptotically tight with respect to this construction, and is exact when $k$ is large and divisible by $b$. We then refine the argument using a circular ordering, obtaining the sharp full-support bound when $k\equiv -1\pmod b$. In the case $b=3$, we prove the exact general result when $k\equiv -1\pmod 3$ and $k$ is sufficiently large. The problem is motivated by the Daniely--Shalev-Shwartz dimension and the pseudocube formulation of a higher-alphabet Sauer-Shelah-Perles lemma.

math.CO

Remarks on the disproof of the unit distance conjecture

We present a short, digested, human-verified version of the recent OpenAI-generated counterexample to the Erd\H{o}s unit distance conjecture, and a sequence of reflections on it. The argument relies crucially on ideas that may, at least in retrospect, be attributed to Ellenberg-Venkatesh, Golod-Shafarevich, and Hajir-Maire-Ramakrishna.

math.CO

Packing arithmetic progressions

Let $\mathcal{F}=\{A_1,A_2,\ldots,A_k\}$ be a collection of finite arithmetic progressions, where each $A_d$ is an initial segment of the set $D_d=\{d,2d,3d,\ldots\}$ of consecutive multiples of a positive integer $d$. Let $m(\mathcal{F})$ denote the minimum length of an interval containing pairwise disjoint \emph{shifted} copies of all members of the family $\mathcal{F}$. We study this parameter in the following two cases: for a fixed positive integer $n$, (1) each progression in $\mathcal{F}$ has the form $A_d=D_d\cap\{1,2,\ldots,n\}$, and (2) all progressions $A_d$ of $\mathcal{F}$ have the same size $n$, that is, $A_d=D_d\cap \{1,2,\ldots, nd\}$. We in particular derive the following asymptotic estimates. In case (1), when $k=n$, we get $m(\mathcal{F})=\Theta(n^{3/2}/\ln n)$. In case (2), when $k=n$, we get $m(\mathcal{F})=\Theta(n^3/\ln n)$, while if $k>k_0(n)$, then $m(\mathcal{F}) < 3kn$. In both cases we additionally determine $m(\mathcal{F})$ asymptotically or settle its order of magnitude for all $k<n$.

math.CO

Aggregating maximal cliques in real-world graphs

Maximal clique enumeration is a fundamental graph mining task, but its utility is often limited by computational intractability and highly redundant output. To address these challenges, we introduce \emph{$\rho$-dense aggregators}, a novel approach that succinctly captures maximal clique structure. Instead of listing all cliques, we identify a small collection of clusters with edge density at least $\rho$ that collectively contain every maximal clique. In contrast to maximal clique enumeration, we prove that for all $\rho < 1$, every graph admits a $\rho$-dense aggregator of \emph{sub-exponential} size, $n^{O(\log_{1/\rho}n)}$, and provide an algorithm achieving this bound. For graphs with bounded degeneracy, a typical characteristic of real-world networks, our algorithm runs in near-linear time and produces near-linear size aggregators. We also establish a matching lower bound on aggregator size, proving our results are essentially tight. In an empirical evaluation on real-world networks, we demonstrate significant practical benefits for the use of aggregators: our algorithm is consistently faster than the state-of-the-art clique enumeration algorithm, with median speedups over $6\times$ for $\rho=0.1$ (and over $300\times$ in an extreme case), while delivering a much more concise structural summary.

cs.DS

Induced matching treewidth and tree-independence number, revisited

We study two graph parameters defined via tree decompositions: tree-independence number and induced matching treewidth. Both parameters are defined similarly as treewidth, but with respect to different measures of a tree decomposition $\mathcal{T}$ of a graph $G$: for tree-independence number, the measure is the maximum size of an independent set in $G$ included in some bag of $\mathcal{T}$, while for the induced matching treewidth, the measure is the maximum size of an induced matching in $G$ such that some bag of $\mathcal{T}$ contains at least one endpoint of every edge of the matching. While the induced matching treewidth of any graph is bounded from above by its tree-independence number, the family of complete bipartite graphs shows that small induced matching treewidth does not imply small tree-independence number. On the other hand, Abrishami, Bria\'nski, Czy\.zewska, McCarty, Milani\v{c}, Rz\k{a}\.zewski, and Walczak~[SIAM Journal on Discrete Mathematics, 2025] showed that, if a fixed biclique $K_{t,t}$ is excluded as an induced subgraph, then the tree-independence number is bounded from above by some function of the induced matching treewidth. The function resulting from their proof is exponential even for fixed $t$, as it relies on multiple applications of Ramsey's theorem. In this note we show, using the K\"ov\'ari-S\'os-Tur\'an theorem, that for any class of $K_{t,t}$-free graphs, the two parameters are in fact polynomially related.

cs.DM

Random Cayley graphs and random sumsets

We prove that any finite abelian group $G$ contains a collection of not too many subsets with a special structure, so that for every subset $A$ of $G$ with a small doubling, there is a member $F$ of the collection that is fully contained in the sumset $A+A$ and is not much smaller than it. Using this result we obtain improved bounds for the problem of estimating the typical independence number of sparse random Cayley or Cayley-sum graphs, and for the problem of estimating the smallest size of a subset of $G$ which is not a sumset. We also obtain tight bounds for the typical maximum length of an arithmetic progression in the sumset of a sparse random subset of $G$.

math.CO

Distinct Directions and Distinct Distances in $\mathbb{R}^d$

We show that there exists an absolute positive constant $b (\geq \frac{1}{48})$ so that any set of $n$ points in $\mathbb{R}^d$ that is $d$-dimensional determines at least $bdn$ lines with pairwise distinct directions. As a consequence we prove that there are $d$-dimensional real norms $\|\cdot\|$ so that every set of $n>n_0(d)$ points that is $d$-dimensional determines at least $(bd-o(1))n$ distinct distances with respect to $\|\cdot \|$.

math.CO

Sums along the edges of bounded degree graphs

Let $G$ be a graph on $n$ vertices and $(H,+)$ be an abelian group. What is the minimum size ${\sf S}_H(G)$ of the set of all sums $A(u)+A(v)$ over all injections $A:V(G)\to H$? In 2012, the first author, Angel, the second author, and Lubetzky proved that, for expander graphs and $H=\mathbb{Z}$, this minimum is at least $\Omega(\log n)$, and this bound is tight -- there exists a regular expander $G$ with ${\sf S}_{\mathbb{Z}}(G)=O(\log n)$. We prove that, for every constant $d\geq 3$, the random $d$-regular graph $\mathcal{G}_{n,d}$ has significantly larger sum-sets: with high probability, for every abelian group $H$, ${\sf S}_H(\mathcal{G}_{n,d})=\Omega(n^{1-2/d})$. In particular, this proves that, for every $\varepsilon>0$, there exists a regular graph with $O(n)$ edges and with sum-sets of size at least $n^{1-\varepsilon}$, for all abelian groups. The bound ${\sf S}_H(\mathcal{G}_{n,d})=\Omega(n^{1-2/d})$ is tight up to a polylogarithmic factor: We show that, for every $3\leq d\leq \ln n/ \ln \ln n$, there exists an abelian group $H$ such that, for every graph $G$ on $n$ vertices with maximum degree at most $d$, ${\sf S}_H(G) \leq n^{1-2/d}(\log n)^{O(1)}$. We also prove that, for $d\gg\ln^2 n$, with high probability, for every abelian group $H$, ${\sf S}_H(\mathcal{G}_{n,d})=n(1-o(1))$ and determine the second-order term, up to a polylogarithmic factor.

math.CO

New Approximation Guarantees for The Inventory Staggering Problem

Since its inception in the mid-60s, the inventory staggering problem has been explored and exploited in a wide range of application domains, such as production planning, stock control systems, warehousing, and aerospace/defense logistics. However, even with a rich history of academic focus, we are still very much in the dark when it comes to cornerstone computational questions around inventory staggering and to related structural characterizations, with our methodological toolbox being severely under-stocked. The central contribution of this paper consists in devising a host of algorithmic techniques and analytical ideas -- some being entirely novel and some leveraging well-studied concepts in combinatorics and number theory -- for surpassing essentially all known approximation guarantees for the inventory staggering problem. In particular, our work demonstrates that numerous structural properties open the door for designing polynomial-time approximation schemes, including polynomially-bounded cycle lengths, constantly-many distinct time intervals, so-called nested instances, and pairwise coprime settings. These findings offer substantial improvements over currently available constant-factor approximations and resolve outstanding open questions in their respective contexts. In parallel, we develop new theory around a number of yet-uncharted questions, related to the sampling complexity of peak inventory estimation as well as to the plausibility of groupwise synchronization. Interestingly, we establish the global nature of inventory staggering, proving that there are $n$-item instances where, for every subset of roughly $\sqrt{n}$ items, no policy improves on the worst-possible one by a factor greater than $1+\epsilon$, whereas for the entire instance, there exists a policy that outperforms the worst-possible one by a factor of nearly $2$, which is optimal.

cs.DS

The spanning tree spectrum: improved bounds and simple proofs

The number of spanning trees of a graph $G$, denoted $\tau(G)$, is a well studied graph parameter with numerous connections to other areas of mathematics. In a recent remarkable paper, answering a question of Sedl\'a\v{c}ek from 1969, Chan, Kontorovich and Pak showed that $\tau(G)$ takes at least $1.1103^n$ different values across simple (and planar) $n$-vertex graphs $G$, for large enough $n$. We give a very short, purely combinatorial proof that at least $1.55^n$ values are attained. We also prove that exponential growth can be achieved with regular graphs, determining the growth rate in another problem first raised by Sedl\'a\v{c}ek in the late 1960's. We further show that the following modular dual version of the result holds. For any integer $N$ and any $u < N$ there exists a planar graph on $O(\log N)$ vertices whose number of spanning trees is $u$ modulo $N$.

math.CO

Circular sorting

We determine the maximal number of steps required to sort $n$ labeled points on a circle by adjacent swaps. Lower bounds for sorting by all swaps, not necessarily adjacent, are given as well.

math.CO

Hitting k primes by dice rolls

Let $S=(d_1,d_2,d_3, \ldots )$ be an infinite sequence of rolls of independent fair dice. For an integer $k \geq 1$, let $L_k=L_k(S)$ be the smallest $i$ so that there are $k$ integers $j \leq i$ for which $\sum_{t=1}^j d_t$ is a prime. Therefore, $L_k$ is the random variable whose value is the number of dice rolls required until the accumulated sum equals a prime $k$ times. It is known that the expected value of $L_1$ is close to $2.43$. Here we show that for large $k$, the expected value of $L_k$ is $(1+o(1)) k\log_e k$, where the $o(1)$-term tends to zero as $k$ tends to infinity. We also include some computational results about the distribution of $L_k$ for $k \leq 100$.

math.PR

A Bicriterion Concentration Inequality and Prophet Inequalities for $k$-Fold Matroid Unions

We investigate prophet inequalities with competitive ratios approaching $1$, seeking to generalize $k$-uniform matroids. We first show that large girth does not suffice: for all $k$, there exists a matroid of girth $\geq k$ and a prophet inequality instance on that matroid whose optimal competitive ratio is $\frac{1}{2}$. Next, we show $k$-fold matroid unions do suffice: we provide a prophet inequality with competitive ratio $1-O(\sqrt{\frac{\log k}{k}})$ for any $k$-fold matroid union. Our prophet inequality follows from an online contention resolution scheme. The key technical ingredient in our online contention resolution scheme is a novel bicriterion concentration inequality for arbitrary monotone $1$-Lipschitz functions over independent items which may be of independent interest. Applied to our particular setting, our bicriterion concentration inequality yields "Chernoff-strength" concentration for a $1$-Lipschitz function that is not (approximately) self-bounding.

cs.DS

On hypercube statistics

Let $d \geq 1$ and $s \leq 2^d$ be nonnegative integers. For a subset $A$ of vertices of the hypercube $Q_n$ and $n\geq d$, let $\lambda(n,d,s,A)$ denote the fraction of subcubes $Q_d$ of $Q_n$ that contain exactly $s$ vertices of $A$. Let $\lambda(n,d,s)$ denote the maximum possible value of $\lambda(n,d,s,A)$ as $A$ ranges over all subsets of vertices of $Q_n$, and let $\lambda(d,s)$ denote the limit of this quantity as $n$ tends to infinity. We prove several lower and upper bounds on $\lambda(d,s)$, showing that for all admissible values of $d$ and $s$ it is larger than $0.28$. We also show that the values of $s=s(d)$ such that $\lambda(d,s)=1$ are exactly $\{0,2^{d-1},2^d\}$. In addition we prove that if $0<s< d/8$, then $\lambda(d, s) \leq 1 - \Omega(1/s)$, and that if $s$ is divisible by a power of $2$ which is $\Omega(s)$ then $\lambda(d,s) \geq 1-O(1/s)$. We suspect that $\lambda(d,1)=(1+o(1))/e$ where the $o(1)$-term tends to $0$ as $d$ tends to infinity, but this remains open, as does the problem of obtaining tight bounds for essentially all other quantities $\lambda(d,s)$.

math.CO

Maximum shattering

A family $\mathcal{F}$ of subsets of $[n]=\{1,2,\ldots,n\}$ shatters a set $A \subseteq [n]$ if for every $A' \subseteq A$ there is an $F \in \mathcal{F}$ such that $F \cap A=A'$. We develop a framework to analyze $f(n,k,d)$, the maximum possible number of subsets of $[n]$ of size $d$ that can be shattered by a family of size $k$. Among other results, we determine $f(n,k,d)$ exactly for $d \leq 2$ and show that if $d$ and $n$ grow, with both $d$ and $n-d$ tending to infinity, then, for any $k$ satisfying $2^d \leq k \leq (1+o(1))2^d$, we have $f(n,k,d)=(1+o(1))c\binom{n}{d}$, where $c$, roughly $0.289$, is the probability that a large square matrix over $\mathbb{F}_2$ is invertible. This latter result extends work of Das and M\'esz\'aros. As an application, we improve bounds for the existence of covering arrays for certain alphabet sizes.

math.CO

Optimal Preprocessing for Answering On-Line Product Queries

We examine the amount of preprocessing needed for answering certain on-line queries as fast as possible. We start with the following basic problem. Suppose we are given a semigroup $(S,\circ )$. Let $s_1 ,\ldots, s_n$ be elements of $S$. We want to answer on-line queries of the form, ``What is the product $s_i \circ s_{i+1} \circ \cdots \circ s_{j-1} \circ s_j$?'' for any given $1\le i\le j\le n$. We show that a preprocessing of $\Theta(n \lambda (k,n))$ time and space is both necessary and sufficient to answer each such query in at most $k$ steps, for any fixed $k$. The function $\lambda (k,\cdot)$ is the inverse of a certain function at the $\lfloor {k/2}\rfloor$-th level of the primitive recursive hierarchy. In case linear preprocessing is desired, we show that one can answer each such query in $O( \alpha (n))$ steps and that this is best possible. The function $\alpha (n)$ is the inverse Ackermann function. We also consider the following extended problem. Let $T$ be a tree with an element of $S$ associated with each of its vertices. We want to answer on-line queries of the form, ``What is the product of the elements associated with the vertices along the path from $u$ to $v$?'' for any pair of vertices $u$ and $v$ in $T$. We derive results that are similar to the above, for the preprocessing needed for answering such queries. All our sequential preprocessing algorithms can be parallelized efficiently to give optimal parallel algorithms which run in $O(\log n)$ time on a CREW PRAM. These parallel algorithms are optimal in both running time and total number of operations. Our algorithms, especially for the semigroup of the real numbers with the minimum or maximum operations, have various applications in certain graph algorithms, in the utilization of communication networks and in Database retrieval.

cs.DS