Searcharxiv⌕ Search

arXiv · 2609.38053

VADER: Filtered Vector Search with Declarative Recall

Abstract

Approximate filtered vector search (FVS), a core operation in many data management tasks that combine structured data with vector embeddings, exhibits increased complexity due to the characteristics of filtering predicates. Each predicate is defined by selectivity (i.e., the fraction of vectors that satisfy the predicate) and correlation (i.e., the relationship between the filter and the vector space), which can significantly affect search difficulty even for the same query vector. This poses a key challenge for users aiming to integrate vector search with structured data, as efficient execution often requires extensive manual tuning of algorithm parameters. In this paper, we present VADER, the first approach that eliminates hyperparameter tuning by introducing declarative recall for approximate filtered vector search. With declarative recall, users specify a desired recall target, and VADER executes FVS queries to meet this target without requiring manual configuration. VADER achieves this by employing a filter-aware recall predictor that generalizes across varying selectivities and correlations without explicit tuning, and by performing early termination once the predicted recall reaches the user-defined target. Through extensive experimental evaluation, we show that VADER achieves near-optimal early termination, while providing significant speedups of up to 53% faster and improved result quality of 28% compared to the best-performing baseline.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Manos Chatzakis, Duo Lu, Helena Caminal, Yannis Chronis, Fatma Özcan, Yannis Papakonstantinou, Timofey Asyrkin, Sebastian Infante Murguia, Itai Rosenblatt, Themis Palpanas. 2026-09-29. VADER: Filtered Vector Search with Declarative Recall. https://arxiv.org/abs/2609.38053

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

X-DigCheck: Co-Evolving Application Profiles and Knowledge Graphs, Demonstrated on the RTI Documentation of Rupe Magna

We demonstrate X-DigCheck, a domain-independent environment for building and maintaining application profiles as they co-evolve with the data they describe. Profiles developed against a fixed ontology quickly drift from the schema they were meant to capture. X-DigCheck treats profile construction as a continuous ontology-data co-evolution loop: data are lifted into RDF against the profile, checked through competency questions and SHACL, and the resulting reports jointly drive revisions of the ontology, mappings, constraints, and graph. The loop is agnostic to the domain and to the pipeline that produces the graph. We validate and demonstrate the tool in the cultural heritage domain, on the construction of RupeMagna-RTI, the first Reflectance Transformation Imaging (RTI) specialisation of the Cultural Heritage Survey ODP (CHS-ODP), aligned with CIDOC-CRM/CRMdig, ArCo, CHAD-KG, and Getty AAT, with semRTI as the lifting pipeline of this use case. The demonstration lets visitors run one full turn of the loop -on the shipped Rupe Magna (Grosio, Italy) RTI survey, or on a profile and graph of their own -executing the competency-question and SHACL checks live and reading the bidirectional coverage report that flags modelling gaps and stale assumptions. The result is a portable co-evolution environment for profile engineering, together with a reusable RTI application profile produced through it. A screencast of the demonstration is available at https://zenodo.org/records/22210609.

cs.DB↗

Transformations for Evolving Property Graph Schemas

Property graph databases are widely used to represent complex and evolving data; yet, systematic support for property graph schema evolution remains limited. In practice, schema transformations are typically defined manually, coupled to specific application contexts, and are difficult to reuse across schemas or evolution scenarios. We present GRAFT, a logic-based framework that models prop- erty graph schema evolution as reusable, order-constrained meta- transformations derived from atomic edits. Schema evolution is formulated as exploration of a finite meta-graph with schemas as nodes and grounded meta-transformations as edges. To ensure tractability, GRAFT combines similarity-guided search and pruning, guaranteeing duplication-freeness, termination and correctness. An experimental evaluation on four benchmark and real-world property graph schema evolution scenarios shows that GRAFT effi- ciently computes high-quality schema transformation sequences. Using greedy exploration, GRAFT reaches the exact target schema on most datasets, producing stable transformation sequences while keeping runtimes low. A qualitative study on both real-world and a synthetic large-scale dataset further shows the quality and robust- ness of the obtained reusable meta-transformations.

cs.DB↗

SAIVE: Selecting AI Valuable Entities

Data lakes store large amounts of telemetry, with logs from network sensors, hosts, and applications containing possibly hundreds of fields for every event. Large enterprises are then left with data lakes that cannot be analyzed efficiently with AI. Aggregate analysis looks at persistent shifts in behavior over time. Many of the fields and columns in data lakes are not useful as they do not contain information that is sufficiently diverse or concentrated to support AI analysis. SAIVE is a simple method for examining a few rows in a large table and applies a histogram of histograms filtering criterion to select the fields that for AI analysis is more likely to yield useful results. This paper provides a principled foundation for the SAIVE heuristics by assuming of a Zipf-Mandelbrot power-law distribution of the underlying data. Constraining the Zipf-Mandelbrot exponent alpha to a reasonable range provides a a practical, cheap, expert-free filter for selecting AI valuable entities in large data sets.

cs.DB↗