Searcharxiv⌕ Search

arXiv · 2610.04520

A Factor-5 Conversion to Internal Collage Systems

Abstract

A collage system is a grammar-based compression model that extends straight-line programs with repetition and substring truncation. Internal collage systems additionally require every nonterminal to be structurally reachable from the start symbol. Migita, Uehata, and I (CPM 2026) showed that any collage system of size $m$ can be converted into an internal one generating the same string with size at most $9m$, and left the improvement of this constant as an open problem. We show that a simple refinement of their top-down conversion reduces the bound to $5m$. The proof combines three elementary ideas: canonicalization of generated truncations that remain aligned with a target endpoint; a repetition decomposition that absorbs every complete copy of the repetition base into a maximal core; and monotonicity of structural reachability, which prevents an input truncation rule from both becoming structurally reachable and later acting as a hidden target at which endpoint alignment is lost. The conversion runs in deterministic $O(m^2)$ worst-case time in the stated unit-cost random-access machine model. Consequently, for every string, the minimum size of an internal collage system is at most five times the minimum size of a general collage system.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Simone Faro. 2026-10-03. A Factor-5 Conversion to Internal Collage Systems. https://arxiv.org/abs/2610.04520

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Non-adaptive Bellman-Ford: Yen's improvement is optimal

The Bellman-Ford algorithm for single-source shortest paths repeatedly updates tentative distances in an operation called {relaxing an edge}. In several important applications a {non-adaptive} (oblivious) implementation is preferred, which means fixing the entire sequence of relaxations upfront, independently of the edge-weights. The original implementation of the algorithm performs, in a dense graph on $n$ vertices, $(1+o(1))n^3 $ relaxations. An improvement by Yen from 1970 reduces the number of relaxations by a factor of two. We show that no further constant-factor improvements are possible, and every {non-adaptive deterministic} algorithm based on relaxations must perform $(\frac{1}{2} - o(1))n^3$ steps. This improves an earlier lower bound of Eppstein of $(\frac{1}{6} - o(1))n^3$. Given that a {non-adaptive randomized} variant of Bellman-Ford with at most $(\frac{1}{3} + o(1))n^3$ relaxations (with high probability) is known, our result implies a strict separation between deterministic and randomized strategies, answering an open question of Eppstein. We also address the complexity of finding {short} relaxation sequences for a given input graph on $n$ vertices, answering a question of Eppstein. We show that the problem is co-NP-hard, and moreover essentially inapproximable: While an $n$-approximation is easily obtained, for every $ε> 0$, no polynomial-time $n^{1-ε}$-approximation exists, unless P = NP. We further show that {deciding} whether a given relaxation sequence is valid is co-NP-complete, even when the input is the complete graph.

cs.DS↗

Exponential Quantum Space Advantage for Approximating Max-$k$SAT in the Streaming Setting

In this paper, we give a one-pass quantum streaming algorithm for Max-$k$SAT that uses $\operatorname{polylog}(n)$ space and achieves a $0.7426$-approximation on instances with $n$ variables and $\operatorname{poly}(n)$ clauses. In contrast, prior work by Chou, Golovnev, and Velusamy (FOCS 2020) implies that achieving an approximation ratio better than $\sqrt{2}/2 \approx 0.7071$ for Max-$k$SAT requires $Ω(\sqrt{n})$ space for any classical streaming algorithm. Therefore, it yields an exponential quantum space advantage for Max-$k$SAT in the streaming setting. Combining with the known results, it gives a complete classification of quantum space advantages for all Boolean Max-2CSPs.

cs.DS↗

Solving Stackelberg Vertex Cover on trees using split and join

The Stackelberg Vertex Cover problem is a bilevel optimization problem with two players on a graph G = ($F \cup P$, E) where each vertex from F has a weight and the first player selects a price for each vertex in P . Afterwards, the second player finds a minimum weight vertex cover X and the first player receives the set price for each vertex from $X \cap P$ . The goal is to maximize the revenue of the first player. This problem was recently shown to be NP-complete for bipartite graphs while being solvable in linear time on paths. We present four new algorithms for solving Stackelberg Vertex Cover on certain kinds of graphs: (1) a pseudo-polynomial algorithm working on general trees when all weights are integer with a runtime linear in the number of vertices and cubic in the maximum weight (2) a generalization of (1) for bipartite graphs with integer weights and a tree decomposition that is FPT in the maximum weight and the treewidth, (3) a strongly polynomial algorithm for rooted trees having the property that the least common ancestor of any two vertices from P is again in P (this case includes paths); and (4) an FPT-algorithm for trees, where the parameter is the maximum number P-vertices $v_i$ that an F-vertex u can reach while using no other P -vertices. These algorithms are based on a lemma that allows us to split instances at a vertex u into multiple sub-instances, which follows from LP duality and integrality of the vertex cover LP on bipartite graphs. The lemma requires that the minimum vertex covers of the sub-instances agree on u (either all include u or all don't). For this we introduce the concept of commitments. We show that the Stackelberg Vertex Cover problem with commitments is weakly NP-complete. An open question is the non-bipartite case as there is an explicit counterexample showing that the split-and-join technique does not work.

cs.DS↗