Searcharxiv⌕ Search

arXiv · 2609.36512

SAIVE: Selecting AI Valuable Entities

Abstract

Data lakes store large amounts of telemetry, with logs from network sensors, hosts, and applications containing possibly hundreds of fields for every event. Large enterprises are then left with data lakes that cannot be analyzed efficiently with AI. Aggregate analysis looks at persistent shifts in behavior over time. Many of the fields and columns in data lakes are not useful as they do not contain information that is sufficiently diverse or concentrated to support AI analysis. SAIVE is a simple method for examining a few rows in a large table and applies a histogram of histograms filtering criterion to select the fields that for AI analysis is more likely to yield useful results. This paper provides a principled foundation for the SAIVE heuristics by assuming of a Zipf-Mandelbrot power-law distribution of the underlying data. Constraining the Zipf-Mandelbrot exponent alpha to a reasonable range provides a a practical, cheap, expert-free filter for selecting AI valuable entities in large data sets.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Inna Voloshchuk, Hayden Jananthan, Jeremy Kepner. 2026-09-29. SAIVE: Selecting AI Valuable Entities. https://arxiv.org/abs/2609.36512

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Epistemic Typing as a PostgreSQL Table Access Method: Adversarial Conflict Resolution Under Confidence Forgery and Sybil Coordination

We describe KNDB, a PostgreSQL 18 table access method (TAM) that types every row with an engine-assigned epistemic kind (MEASURED, INFERRED, or DERIVED) and resolves per-slot conflicts inside every write-time heapam callback. Rows land as ordinary heap tuples; seven of the 44 TAM callbacks are overridden (tuple_insert, multi_insert, tuple_update, tuple_delete, tuple_insert_speculative, tuple_complete_speculative, relation_toast_am), the other 37 delegate to heap; we provide a completeness argument over the interface as a paper artefact. This paper reports the engineering behind that decision and the adversarial evaluation that motivated it. On a confidence-forgery workload where an attacker asserts INFERRED writes with confidence in [0.95,1.0] against honest MEASURED writes with confidence in [0.5,0.9], KNDB beats a confidence-only baseline by 63 percentage points on the Book-Author fusion dataset and 92.7 points on the Zheng crowdsourcing dataset. Both wins are proven load-bearing on the kind axis by a source-rebuild disable-and-test in which the lattice is neutralised and the win vanishes. Against four truth-discovery baselines (TruthFinder, CRH, CATD, ACCU) reimplemented from the original equations and validated to within 0.3 percentage points of the published numbers, KNDB is competitive below a per-dataset density-saturation cell and dominant at or above it. We formalise the cell as k* ~ rho_alg * h_top, where h_top is per-slot top honest surface-form support, and validate the prediction within +/-20% on Book-Author and +/-30% on Zheng. Because the kind axis is assigned by the engine from independent metadata and cannot be forged at write time, KNDB's k* is unbounded. The paper is honest about where KNDB loses: CRH and ACCU outperform KNDB below saturation on Zheng, and KNDB scores zero on three temporal knowledge-editing benchmarks whose ground truth is last-writer-wins.

cs.DB↗

WeaveData: A Multimodal Data Analysis System with Self-Critiquing and Self-Evolving LLM Plans

Multimodal data analysis, which answers questions over relational tables, text, and images, has attracted growing attention in the data management community. Large language models (LLMs) enable such analysis in natural language by generating analysis plans over relational and semantic operators. However, LLM-generated plans are error-prone: a plan may silently compute something other than what was asked, fail during execution, or return a result that misses the question. This paper presents WeaveData, a multimodal data analysis system with self-critiquing and self-evolving LLM plans. First, WeaveData generates a typed logical plan for each question and critiques it step by step before execution, and it checks the executed result against the question afterwards. Second, WeaveData evolves a plan that fails or misses the question: it diagnoses the failure with the actual data, reuses the results that remain valid, and accumulates planning experience for later questions. Third, WeaveData grounds planning in a metadata knowledge graph of all modalities, clarifies ambiguous questions with the user, and backs every model judgment with evidence in an interactive notebook. We demonstrate WeaveData on two public multimodal datasets.

cs.DB↗

Transformations for Evolving Property Graph Schemas

Property graph databases are widely used to represent complex and evolving data; yet, systematic support for property graph schema evolution remains limited. In practice, schema transformations are typically defined manually, coupled to specific application contexts, and are difficult to reuse across schemas or evolution scenarios. We present GRAFT, a logic-based framework that models prop- erty graph schema evolution as reusable, order-constrained meta- transformations derived from atomic edits. Schema evolution is formulated as exploration of a finite meta-graph with schemas as nodes and grounded meta-transformations as edges. To ensure tractability, GRAFT combines similarity-guided search and pruning, guaranteeing duplication-freeness, termination and correctness. An experimental evaluation on four benchmark and real-world property graph schema evolution scenarios shows that GRAFT effi- ciently computes high-quality schema transformation sequences. Using greedy exploration, GRAFT reaches the exact target schema on most datasets, producing stable transformation sequences while keeping runtimes low. A qualitative study on both real-world and a synthetic large-scale dataset further shows the quality and robust- ness of the obtained reusable meta-transformations.

cs.DB↗