SearcharxivSearch

arXiv subjects

Kevin Kristensen

Publications and source records attributed to Kevin Kristensen.

3 recordsLinked to original sources

Optimizing Bloom Filters on Modern GPUs

Bloom filters are a fundamental data structure for approximate membership queries in applications ranging from analytics and databases to genomics. Deployed as prefilters, they eliminate irrelevant data before expensive downstream processing. As data-processing pipelines move onto GPUs, filtering must remain GPU-resident and keep pace with other stages. Although Bloom filters have been extensively optimized for CPUs, few implementations target GPUs, where fixed SIMD layouts map poorly to SIMT hardware and leave performance potential on the table. We present an architecture-aware GPU Bloom filter with tunable vectorization for performance portability across workloads, memory regimes, and GPU architectures. On NVIDIA B200, it sustains over $92\%$ of the measured random-access bound. At comparable false-positive rates, it outperforms the state-of-the-art GPU baseline by $15.4\times$ for lookup and $11.35\times$ for construction. These gains bring accurate Bloom filters to GPU-scale throughput previously reserved for high-error variants. The implementation is openly available in NVIDIA's cuCollections library: https://github.com/NVIDIA/cuCollections.

cs.DC

One Join Order Does Not Fit All: Reducing Intermediate Results with Per-Split Query Plans

Minimizing intermediate results is critical for efficient multi-join query processing. Although the seminal Yannakakis algorithm offers strong guarantees for acyclic queries, cyclic queries remain an open challenge. In this paper, we propose SplitJoin, a framework that introduces split as a first-class query operator. By partitioning input tables into heavy and light parts, SplitJoin allows different data partitions to use distinct query plans, with the goal of reducing intermediate sizes using existing binary join engines. We systematically explore the design space for split-based optimizations, including threshold selection, split strategies, and join ordering after splits. Implemented as a front-end to DuckDB and Umbra, SplitJoin achieves substantial improvements: on DuckDB, SplitJoin completes 43 social network queries (vs. 29 natively), achieving 2.1x faster runtime and 7.9x smaller intermediates on average (up to 13.6x and 74x, respectively); on Umbra, it completes 45 queries (vs. 35), achieving 1.3x speedups and 1.2x smaller intermediates on average (up to 6.1x and 2.1x, respectively).

cs.DB

Rethinking Analytical Processing in the GPU Era

The era of GPU-powered data analytics has arrived. In this paper, we argue that recent advances in hardware (e.g., larger GPU memory, faster interconnect and IO, and declining cost) and software (e.g., composable data systems and mature libraries) have removed the key barriers that have limited the wider adoption of GPU data analytics. We present Sirius, a prototype open-source GPU-native SQL engine that offers drop-in acceleration for diverse data systems. Sirius treats GPU as the primary engine and leverages libraries like libcudf for high-performance relational operators. It provides drop-in acceleration for existing databases by leveraging the standard Substrait query representation, replacing the CPU engine without changing the user-facing interface. Sirius achieves 8.3x and 7.4x better cost efficiency on TPC-H and ClickBench, respectively, when integrated with single-node DuckDB, and delivers up to 12.5x speedup when integrated with Apache Doris distributed engine.

cs.DB