arXiv · 2609.12754
Learning the Lake: Reliable Experience for Adaptive Data Product Discovery
Abstract
Data-product discovery searches a full lake even when workloads revisit related products and regions. Repetition permits contracted search, but similarity cannot justify a route because one omitted asset invalidates a conjunctive product. We study when serving experience can safely reduce this work. Evolving Discovery Memory records source-labelled query--product--region evidence above a fixed regional index. SafeLake separates operational familiarity, which determines how much to search, from independently calibrated product evidence, which determines where to search. The fixed-probe comparison holds the adaptive budget constant between SafeLake and Familiarity-only. On TAT-QA, product steering raises Product Recall by 0.072; ConvFinQA shows no resolved map gain, while the HybridQA sensitivity favors Familiarity-only in Full R@100. Trace-only, missing, and false feedback expose boundaries on map steering, while scope-audit agreement cannot certify the source. Across clean confirmed-feedback streams under the frozen transductive protocol, the formal controller saves 49.5--82.7% of cumulative asset exposure. Experience determines when to contract; reliable evidence determines where to contract.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yixi Zhou, Fan Zhang, Sikun Wang, Yingfan Xu, Haipeng Zhang. 2026-09-11. Learning the Lake: Reliable Experience for Adaptive Data Product Discovery. https://arxiv.org/abs/2609.12754
Cite the original work for its findings. Save a collection to share your selection of sources.