arXiv · 2310.02656
Blend: A Unified Data Discovery System
Abstract
Most research on data discovery has so far focused on improving individual discovery operators such as join, correlation, or union discovery. However, in practice, a combination of these techniques and their corresponding indexes may be necessary to support arbitrary discovery tasks. We propose BLEND, a comprehensive data discovery system that supports existing operators and enables their flexible pipelining. BLEND is based on a set of lower-level operators that serve as fundamental building blocks for more complex and sophisticated user tasks. To reduce the execution runtime of discovery pipelines, we propose a unified index structure and a rule-based optimizer that rewrites SQL statements into low-level operators when possible. We show the superior flexibility and efficiency of our system compared to ad-hoc discovery pipelines and stand-alone solutions.
Explore related subjects
Keep this discovery
Mahdi Esmailoghli, Christoph Schnell, Renée J. Miller, Ziawasch Abedjan. 2023-10-04. Blend: A Unified Data Discovery System. https://arxiv.org/abs/2310.02656
Cite the original work for its findings. Save a collection to share your selection of sources.