SearcharxivSearch

arXiv subjects

Keita Iwabuchi

Publications and source records attributed to Keita Iwabuchi.

9 recordsLinked to original sources

SOLANET: Distributed Neighbor Graph Construction on GPU-Accelerated Systems

Neighbor graphs capture relationships among data points and are widely used in data analytics and AI workloads. Many studies have explored approximate construction methods for single-node systems, including GPUs. However, extending this to distributed systems for larger data and further acceleration remains challenging due to irregular computation patterns. We present SOLANET, a GPU-accelerated distributed neighbor graph construction toolkit. SOLANET first constructs local graphs on each GPU after data partitioning and then refines them via approximate nearest neighbor (ANN) searches over remote graphs pulled from other GPUs using MPI one-sided operations. SOLANET also provides a lock-free single-GPU neighbor graph construction algorithm for AMD GPUs. Our single-GPU implementation outperforms a state-of-the-art GPU-based approximate neighbor graph construction implementation across multiple datasets on a single MI300A APU. Furthermore, SOLANET demonstrates 11X speedup from 32 to 512 APUs for 1 billion data points and 6.9x speedup from 64 to 512 APUs for 2 billion points.

cs.DC

Better Learning-Augmented Spanning Tree Algorithms via Metric Forest Completion

We present improved learning-augmented algorithms for finding an approximate minimum spanning tree (MST) for points in an arbitrary metric space. Our work follows a recent framework called metric forest completion (MFC), where the learned input is a forest that must be given additional edges to form a full spanning tree. Veldt et al. (2025) showed that optimally completing the forest takes $Ω(n^2)$ time, but designed a 2.62-approximation for MFC with subquadratic complexity. The same method is a $(2γ+ 1)$-approximation for the original MST problem, where $γ\geq 1$ is a quality parameter for the initial forest. We introduce a generalized method that interpolates between this prior algorithm and an optimal $Ω(n^2)$-time MFC algorithm. Our approach considers only edges incident to a growing number of strategically chosen ``representative'' points. One corollary of our analysis is to improve the approximation factor of the previous algorithm from 2.62 for MFC and $(2γ+1)$ for metric MST to 2 and $2γ$ respectively. We prove this is tight for worst-case instances, but we still obtain better instance-specific approximations using our generalized method. We complement our theoretical results with a thorough experimental evaluation.

cs.DS

Multi-year stacking searches for solar system bodies

Digital tracking detects faint solar system bodies by stacking many images along hypothesized orbits, revealing objects that are undetectable in every individual exposure. Previous searches have been restricted to small areas and short time baselines. We present a general framework to quantify both sensitivity and computational requirements for digital tracking of nonlinear motion across the full sky over multi-year baselines. We start from matched-filter stacking and derive how signal-to-noise ratio (SNR) degrades with trial orbit mismatch, which leads to a metric tensor on orbital parameter space. The metric defines local Euclidean coordinates in which SNR loss is isotropic, and a covariant density that specifies the exact number of trial orbits needed for a chosen SNR tolerance. We validate the approach with Zwicky Transient Facility (ZTF) data, recovering known objects in blind searches that stack thousands of images over six years along billions of trial orbits. We quantify ZTF's sensitivity to populations beyond 5 au and show that stacking reaches most of the remaining Planet 9 parameter space. The computational demands of all-sky, multi-year tracking are extreme, but we demonstrate that time segmentation and image blurring greatly reduce orbit density at modest sensitivity cost. Stacking effectively boosts medium-aperture surveys to the Rubin Observatory single-exposure depth across the northern sky. Digital tracking in dense Rubin observations of a 10 sq. deg field is tractable and could detect trans-Neptunian objects to 27th magnitude in a single night, with deep drilling fields reaching fainter still.

astro-ph.EP

Survey-Wide Asteroid Discovery with a High-Performance Computing Enabled Non-Linear Digital Tracking Framework

Modern astronomical surveys detect asteroids by linking together their appearances across multiple images taken over time. This approach faces limitations in detecting faint asteroids and handling the computational complexity of trajectory linking. We present a novel method that adapts ``digital tracking" - traditionally used for short-term linear asteroid motion across images - to work with large-scale synoptic surveys such as the Vera Rubin Observatory Legacy Survey of Space and Time (Rubin/LSST). Our approach combines hundreds of sparse observations of individual asteroids across their non-linear orbital paths to enhance detection sensitivity by several magnitudes. To address the computational challenges of processing massive data sets and dense orbital phase spaces, we developed a specialized high-performance computing architecture. We demonstrate the effectiveness of our method through experiments that take advantage of the extensive computational resources at Lawrence Livermore National Laboratory. This work enables the detection of significantly fainter asteroids in existing and future survey data, potentially increasing the observable asteroid population by orders of magnitude across different orbital families, from near-Earth objects (NEOs) to Kuiper belt objects (KBOs).

astro-ph.EP

Approximate Tree Completion and Learning-Augmented Algorithms for Metric Minimum Spanning Trees

Finding a minimum spanning tree (MST) for $n$ points in an arbitrary metric space is a fundamental primitive for hierarchical clustering and many other ML tasks, but this takes $Ω(n^2)$ time to even approximate. We introduce a framework for metric MSTs that first (1) finds a forest of disconnected components using practical heuristics, and then (2) finds a small weight set of edges to connect disjoint components of the forest into a spanning tree. We prove that optimally solving the second step still takes $Ω(n^2)$ time, but we provide a subquadratic 2.62-approximation algorithm. In the spirit of learning-augmented algorithms, we then show that if the forest found in step (1) overlaps with an optimal MST, we can approximate the original MST problem in subquadratic time, where the approximation factor depends on a measure of overlap. In practice, we find nearly optimal spanning trees for a wide range of metrics, while being orders of magnitude faster than exact algorithms.

cs.DS

A Scalable Gaussian Process Approach to Shear Mapping with MuyGPs

Analysis of cosmic shear is an integral part of understanding structure growth across cosmic time, which in-turn provides us with information about the nature of dark energy. Conventional methods generate \emph{shear maps} from which we can infer the matter distribution in the universe. Current methods (e.g., Kaiser-Squires inversion) for generating these maps, however, are tricky to implement and can introduce bias. Recent alternatives construct a spatial process prior for the lensing potential, which allows for inference of the convergence and shear parameters given lensing shear measurements. Realizing these spatial processes, however, scales cubically in the number of observations - an unacceptable expense as near-term surveys expect billions of correlated measurements. Therefore, we present a linearly-scaling shear map construction alternative using a scalable Gaussian Process (GP) prior called MuyGPs. MuyGPs avoids cubic scaling by conditioning interpolation on only nearest-neighbors and fits hyperparameters using batched leave-one-out cross validation. We use a suite of ray-tracing results from N-body simulations to demonstrate that our method can accurately interpolate shear maps, as well as recover the two-point and higher order correlations. We also show that we can perform these operations at the scale of billions of galaxies on high performance computing platforms.

astro-ph.CO

Metall: A Persistent Memory Allocator For Data-Centric Analytics

Data analytics applications transform raw input data into analytics-specific data structures before performing analytics. Unfortunately, such data ingestion step is often more expensive than analytics. In addition, various types of NVRAM devices are already used in many HPC systems today. Such devices will be useful for storing and reusing data structures beyond a single process life cycle. We developed Metall, a persistent memory allocator built on top of the memory-mapped file mechanism. Metall enables applications to transparently allocate custom C++ data structures into various types of persistent memories. Metall incorporates a concise and high-performance memory management algorithm inspired by Supermalloc and the rich C++ interface developed by Boost.Interprocess library. On a dynamic graph construction workload, Metall achieved up to 11.7x and 48.3x performance improvements over Boost.Interprocess and memkind (PMEM kind), respectively. We also demonstrate Metall's high adaptability by integrating Metall into a graph processing framework, GraphBLAS Template Library. This study's outcomes indicate that Metall will be a strong tool for accelerating future large-scale data analytics by allowing applications to leverage persistent memory efficiently.

cs.DC

TriPoll: Computing Surveys of Triangles in Massive-Scale Temporal Graphs with Metadata

Understanding the higher-order interactions within network data is a key objective of network science. Surveys of metadata triangles (or patterned 3-cycles in metadata-enriched graphs) are often of interest in this pursuit. In this work, we develop TriPoll, a prototype distributed HPC system capable of surveying triangles in massive graphs containing metadata on their edges and vertices. We contrast our approach with much of the prior effort on triangle analysis, which often focuses on simple triangle counting, usually in simple graphs with no metadata. We assess the scalability of TriPoll when surveying triangles involving metadata on real and synthetic graphs with up to hundreds of billions of edges.We utilize communication-reducing optimizations to demonstrate a triangle counting task on a 224 billion edge web graph in approximately half of the time of competing approaches, while additionally supporting metadata-aware capabilities.

cs.DC

UMap: Enabling Application-driven Optimizations for Page Management

Leadership supercomputers feature a diversity of storage, from node-local persistent memory and NVMe SSDs to network-interconnected flash memory and HDD. Memory mapping files on different tiers of storage provides a uniform interface in applications. However, system-wide services like mmap are optimized for generality and lack flexibility for enabling application-specific optimizations. In this work, we present Umap to enable user-space page management that can be easily adapted to access patterns in applications and storage characteristics. Umap uses the userfaultfd mechanism to handle page faults in multi-threaded applications efficiently. By providing a data object abstraction layer, Umap is extensible to support various backing stores. The design of Umap supports dynamic load balancing and I/O decoupling for scalable performance. Umap also uses application hints to improve the selection of caching, prefetching, and eviction policies. We evaluate Umap in five benchmarks and real applications on two systems. Our results show that leveraging application knowledge for page management could substantially improve performance. On average, Umap achieved 1.25 to 2.5 times improvement using the adapted configurations compared to the system service.

cs.DC