SearcharxivSearch

arXiv subjects

Ravi Shetye

Publications and source records attributed to Ravi Shetye.

3 recordsLinked to original sources

Tursio for Credit Unions: Structured Data Search with Automated Context Graphs

Extracting actionable insights from structured databases in regulated industries, such as credit unions, is often hindered by complex schemas, legacy systems, and stringent data governance requirements. We present Tursio, a secure, on-premises, database search platform that enables business users to query enterprise databases using natural language. Tursio automatically infers a context graph -- a schema-level metadata structure that captures join paths, column semantics, and domain annotations -- and uses it to systematically generate accurate query plans through LLM-assisted compilation, grounding, and rewriting. Unlike existing AI/BI tools that require extensive manual context curation, Tursio automates this end-to-end and deploys entirely on-premises. We demonstrate Tursio through realistic scenarios in the credit union domain, and discuss its applicability to other regulated settings.

cs.DB

Scalable Join Inference for Large Context Graphs

Context graphs are essential for modern AI applications including question answering, pattern discovery, and data analysis. Building accurate context graphs from structured databases requires inferring join relationships between entities. Invalid joins introduce ambiguity and duplicate records, compromising graph quality. We present a scalable join inference approach combining statistical pruning with Large Language Model (LLM) reasoning. Unlike purely statistics-based methods, our hybrid approach mimics human semantic understanding while mitigating LLM hallucination through data-driven inference. We first identify primary key candidates and use LLMs for adjudication, then detect inclusion dependencies with the same two-stage process. This statistics-LLM combination scales to large schemas while maintaining accuracy and minimizing false positives. We further leverage the database query history to refine the join inferences over time as the query workloads evolve. Our evaluation on TPC-DS, TPC-H, BIRD-Dev, and production workloads demonstrates that the approach achieves high precision (78-100%) on well-structured schemas, while highlighting the inherent difficulty of join discovery in poorly normalized settings.

cs.DB

Making Databases Searchable with Deep Context

Databases are the most critical assets for enterprises, and yet they remain largely inaccessible to people who make the most important decisions. In this paper, we describe the Tursio search platform that builds an abstraction layer, aka semantic knowledge graph, over the underlying databases to make them searchable in natural language. Tursio infuses large language models (LLMs) into every part of the query processing stack, including data modeling, query compilation, query planning, and result reasoning. This allows Tursio to process natural language queries systematically using techniques from traditional query planning and rewriting, rather than black-box memorization. We describe the architecture of Tursio in detail and present a comprehensive evaluation on production workloads, and synthetic and realistic benchmarks. Our results show that Tursio achieves high accuracy while being efficient and scalable, making databases truly searchable for non-expert users.

cs.DB