Searcharxiv⌕ Search

arXiv · 2609.30956

Attacking Diophantus: Special Cases of Bag Containment

Abstract

Query containment is a fundamental decision problem in database theory: given two queries, determine whether, over all database instances, every answer produced by the first is also produced by the second. For conjunctive queries under set semantics, the problem is understood through the classical homomorphism-based characterisation. Under bag semantics, the interpretation underlying real relational databases, containment becomes a quantitative comparison of answer multiplicities. Despite decades of work, the decidability of bag containment for conjunctive queries remains open. This frontier is fragile: for slightly more expressive classes, bag containment is undecidable, with negative results relying on reductions from variants of Hilbert's 10th problem. This work develops a unified framework for bag containment of conjunctive queries that subsumes two previously studied decidable cases: projection-free and join-on-free containee queries. The framework yields decidability for a broader class, called join-uniform queries, while leaving the containing query arbitrary. This contrasts with techniques that impose restrictions on the containing query. The approach identifies tractable classes based on the internal unification structure of the query whose multiplicities must be bounded. Specifically, it reduces containment to a controlled Diophantine problem. Starting from the containee query, one builds a canonical model generated by all its possible unifications, over which multiplicities admit a finite arithmetic characterisation. Containment is proved equivalent to the non-existence of solutions of a corresponding Diophantine inequality system. Although these problems are undecidable in general, we show that the systems arising from join-uniform containment form a decidable subclass. Thus, the standard source of undecidability for bag containment becomes the core of the decision procedure.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

George Konstantinidis, Xinzhuo Li, Fabio Mogavero. 2026-09-25. Attacking Diophantus: Special Cases of Bag Containment. https://arxiv.org/abs/2609.30956

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Tree Databases

We propose a novel database model whose basic structure is a labeled, directed tree with node identities. Intuitively, the root of the tree is seen as an object (or entity), the non-root nodes as attributes of the object and the semantics of each attribute is represented by the unique path leading from the root to the attribute. We define a tree database to be a set of such trees. The trees of the database can be combined to produce new trees using a set of operations on trees that we define in the paper. The query language of our model offers two types of queries, traversal queries and analytic queries. A query (whether traversal or analytic) is always defined over a tree, which is either a tree in the database or a tree derived from other trees using tree operations. The operations on trees and the query language are both defined using a simple functional algebra whose operations are: restriction of a function, composition of functions, pairing of functions and Cartesian product of sets. A distinctive feature of our model is that traversal queries and analytic queries are both defined within the same formal framework; and in fact, traversal queries serve as the building blocks for analytic queries. This is in sharp contrast to the relational model, where analytic queries are defined outside the relational algebra, in the form of SQL Group-by queries. Therefore our model supports data access and data analysis within the same formal framework. We demonstrate the expressive power of our model by showing: (a) how our model can support inheritance in a seamless manner, (b) how one can define consistent relational databases on top of a tree database - with the tree database playing the role of an underlying semantic layer and (c) how a tree database can be used as a user-friendly interface for accessing and analyzing relational data.

cs.DB↗

Batched Feedback and the Random-Access Wall in Search-Based Graph Construction

Navigable graphs for nearest-neighbor search are built either incrementally, each inserted point searching a graph that mutates as construction proceeds, or in batch over a fixed substrate, which buys determinism and parallelism at a price in build time. We measure that price and find where it comes from. Instrumenting a tuned Vamana and PiPNN and building every system repeatedly in a paired design on one 64-thread machine, we separate build time into work (distance evaluations) and cost per evaluation. Letting the batch builder's substrate mutate in $B$ synchronous blocks recovers the feedback loop of incremental construction while the graph stays a function of (data, seed, parameters, $B$): eight trees with batched feedback match thirty-two frozen trees, and the build does $0.88\times$ Vamana's distance work. Yet it takes $1.55\times$ Vamana's wall-clock, because the mutating substrate costs more per evaluation. Pushing further, we find a wall that no search-based builder crosses. Every one of them, incremental or batched, evaluates distances adaptively, one dependent random access at a time, and runs $20$-$27\times$ below a dense kernel on the same machine at $d \approx 100$, five to six times of it with the data resident in cache. The gap is not a low-dimension artifact: an adaptive evaluation costs $d^{1.04}$ and a blocked one $d^{0.63}$, so the wall grows as $d^{0.4}$ and is twice as high at $d = 960$ as at $d = 128$. PiPNN's order-of-magnitude build advantage is that kernel: it evaluates as many distances per point, but as fixed-in-advance dense blocks. We show the beam cannot be batched after the fact (the useful density of a lockstep block is 2-3%), state the wall as a two-ceiling roofline, and delimit it: it binds whenever the distance is a black box or the evaluation order is data-dependent.

cs.DB↗

MLSkip: Data Skipping for ML Filters via Lightweight Metadata

Database vendors recently released AI functions that can be used in filter predicates. As such functions often rely on costly, black-box ML models, they unveil new data management challenges. Concretely, traditional data skipping techniques for integer and string data fail to be applicable to the new filter type. Indeed, there is no known mechanism for pruning non-qualifying row groups, e.g., when reading files from blob storage. In this work, we initiate the study of data skipping techniques for ML filters. We make the case that Parquet's default min-max metadata is enough to enable pruning. To this end, we draw connections to two lines of research: (i) the recently proposed query language for ML models and (ii) neural network verification. Our preliminary results on ReLU architectures show that on tables from TPC-H and TPC-DS, the average pruning effectiveness for filters of selectivity below 0.1% amounts to 27.4%. Finally, inspired by research on spatial joins, we propose an enhanced metadata structure: a size-bounded 2D convex hull that verification tools can make better use of, increasing the pruning effectiveness to 38.31%, while occupying at most 45 bytes per row group and column pair. We observe an end-to-end speedup of 1.07$\times$ over PyTorch in DuckDB.

cs.DB↗