SearcharxivSearch

arXiv subjects

John Bachan

Publications and source records attributed to John Bachan.

5 recordsLinked to original sources

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives

GPU collective communication is typically optimized for bandwidth, yet many emerging workloads are increasingly limited by latency. Long-context decode-heavy large language model (LLM) inference is a prime example, where serving large models requires multiple GPUs, and many small collectives lie directly on the critical path of token generation. Therefore, even microsecond of overhead can impact performance and cost. In this work, we study how to approach the hardware Speed-of-Light (SoL) lower bound for GPU collectives within a scale-up network. We identify key principles for near-optimal designs, including barrier-free synchronization and efficient use of symmetric memory and multicast. Building on NCCL's device-side API, we develop low-latency interfaces for constructing custom collective kernels and use them to implement new symmetric collectives in NCCL. Microbenchmarks show substantial latency reductions for small and medium messages, reducing overhead to within 7% of the absolute SoL lower bound. When integrated into real applications, these kernels improve inter-token latency and throughput in LLM inference and accelerate cuSOLVERMp, demonstrating benefits for both AI inference and traditional HPC workloads.

cs.DC

GPU-Initiated Networking for NCCL

Modern AI workloads, especially Mixture-of-Experts (MoE) architectures, increasingly demand low-latency, fine-grained GPU-to-GPU communication with device-side control. Traditional GPU communication follows a host-initiated model, where the CPU orchestrates all communication operations - a characteristic of the CUDA runtime. Although robust for collective operations, applications requiring tight integration of computation and communication can benefit from device-initiated communication that eliminates CPU coordination overhead. NCCL 2.28 introduces the Device API with three operation modes: Load/Store Accessible (LSA) for NVLink/PCIe, Multimem for NVLink SHARP, and GPU-Initiated Networking (GIN) for network RDMA. This paper presents the GIN architecture, design, semantics, and highlights its impact on MoE communication. GIN builds on a three-layer architecture: i) NCCL Core host-side APIs for device communicator setup and collective memory window registration; ii) Device-side APIs for remote memory operations callable from CUDA kernels; and iii) A network plugin architecture with dual semantics (GPUDirect Async Kernel-Initiated and Proxy) for broad hardware support. The GPUDirect Async Kernel-Initiated backend leverages DOCA GPUNetIO for direct GPU-to-NIC communication, while the Proxy backend provides equivalent functionality via lock-free GPU-to-CPU queues over standard RDMA networks. We demonstrate GIN's practicality through integration with DeepEP, an MoE communication library. Comprehensive benchmarking shows that GIN provides device-initiated communication within NCCL's unified runtime, combining low-latency operations with NCCL's collective algorithms and production infrastructure.

cs.DC

Flash-X, a multiphysics simulation software instrument

Flash-X is a highly composable multiphysics software system that can be used to simulate physical phenomena in several scientific domains. It derives some of its solvers from FLASH, which was first released in 2000. Flash-X has a new framework that relies on abstractions and asynchronous communications for performance portability across a range of increasingly heterogeneous hardware platforms. Flash-X is meant primarily for solving Eulerian formulations of applications with compressible and/or incompressible reactive flows. It also has a built-in, versatile Lagrangian framework that can be used in many different ways, including implementing tracers, particle-in-cell simulations, and immersed boundary methods.

physics.comp-ph

Adaptive Total Variation Stable Local Timestepping for Conservation Laws

This paper proposes a first-order total variation diminishing (TVD) treatment for coarsening and refining of local timestep size in response to dynamic local variations in wave speeds for nonlinear conservation laws. The algorithm is accompanied with a proof of formal correctness showing that given a sufficiently small minimum timestep the algorithm will produce TVD solution for nonlinear scalar conservation laws. A key feature of the algorithm is its formulation as a discrete event simulation, which allows for easy and efficient parallelization using existing software. Numerical results demonstrate the stability and adaptivity of the method for the shallow water equations. We also introduce a performance model to load balance and explain the observed performance gains. Performance results are presented for a single node on Stampede2's Skylake partition using an optimistic parallel discrete event simulator. Results show the proposed algorithm recovering 59%-77% of the theoretically achievable speed-up with the discrepancies being attributed to the cost of computing the CFL condition and load imbalance.

math.NA

ExaGridPF: A Parallel Power Flow Solver for Transmission and Unbalanced Distribution Systems

This paper investigates parallelization strategies for solving power flow problems in both transmission and unbalanced, three-phase distribution systems by developing a scalable power flow solver, ExaGridPF, which is compatible with existing high-performance computing platforms. Newton-Raphson (NR) and Newton-Krylov (NK) algorithms have been implemented to verify the performance improvement over both standard IEEE test cases and synthesized grid topologies. For three-phase, unbalanced system, we adapt the current injection method (CIM) to model the power flow and utilize SuperLU to parallelize the computing load across multiple threads. The experimental results indicate that more than 5 times speedup ratio can be achieved for synthesized large-scale transmission topologies, and significant efficiency improvements are observed over existing methods for the distribution networks.

cs.CE