SearcharxivSearch

arXiv subjects

Sreevatsa Anantharamu

Publications and source records attributed to Sreevatsa Anantharamu.

9 recordsLinked to original sources

The Multipath Reliable Connection (MRC) Transport

MRC is an open, production-grade transport designed for large-scale AI/ML training over best-effort Ethernet. It extends RoCEv2 with explicit, composable primitives for per-packet multipath and sender-based congestion control, decouples packet delivery from semantic processing, adds multiple new capabilities for accelerated packet-loss recovery and adds resilience against port and path failures. This paper presents MRC and details its core capabilities and mechanisms.

cs.NI

MSCCL++: Rethinking GPU Communication Abstractions for AI Inference

AI applications increasingly run on fast-evolving, heterogeneous hardware to maximize performance, but general-purpose libraries lag in supporting these features. Performance-minded programmers often build custom communication stacks that are fast but error-prone and non-portable. This paper introduces MSCCL++, a design methodology for developing high-performance, portable communication kernels. It provides (1) a low-level, performance-preserving primitive interface that exposes minimal hardware abstractions while hiding the complexities of synchronization and consistency, (2) a higher-level DSL for application developers to implement workload-specific communication algorithms, and (3) a library of efficient algorithms implementing the standard collective API, enabling adoption by users with minimal expertise. Compared to state-of-the-art baselines, MSCCL++ achieves geomean speedups of $1.7\times$ (up to $5.4\times$) for collective communication and $1.2\times$ (up to $1.38\times$) for AI inference workloads. MSCCL++ is in production of multiple AI services provided by Microsoft Azure, and has also been adopted by RCCL, the GPU collective communication library maintained by AMD. MSCCL++ is open source and available at https://github.com/microsoft/mscclpp . Our two years of experience with MSCCL++ suggests that its abstractions are robust, enabling support for new hardware features, such as multimem, within weeks of development.

cs.DC

SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication

RDMA is vital for efficient distributed training across datacenters, but millisecond-scale latencies complicate the design of its reliability layer. We show that depending on long-haul link characteristics, such as drop rate, distance and bandwidth, the widely used Selective Repeat algorithm can be inefficient, warranting alternatives like Erasure Coding. To enable such alternatives on existing hardware, we propose SDR-RDMA, a software-defined reliability stack for RDMA. Its core is a lightweight SDR SDK that extends standard point-to-point RDMA semantics -- fundamental to AI networking stacks -- with a receive buffer bitmap. SDR bitmap enables partial message completion to let applications implement custom reliability schemes tailored to specific deployments, while preserving zero-copy RDMA benefits. By offloading the SDR backend to NVIDIA's Data Path Accelerator (DPA), we achieve line-rate performance, enabling efficient inter-datacenter communication and advancing reliability innovation for inter-datacenter training.

cs.NI

Efficient implementation of the hybridized Raviart-Thomas mixed method by converting flux subspaces into stabilizations

We show how to reduce the computational time of the practical implementation of the Raviart-Thomas mixed method for second-order elliptic problems. The implementation takes advantage of a recent result which states that certain local subspaces of the vector unknown can be eliminated from the equations by transforming them into stabilization functions; see the paper published online in JJIAM on August 10, 2023. We describe in detail the new implementation (in MATLAB and a laptop with Intel(R) Core (TM) i7-8700 processor which has six cores and hyperthreading) and present numerical results showing 10 to 20% reduction in the computational time for the Raviart-Thomas method of index $k$, with $k$ ranging from 1 to 20, applied to a model problem.

math.NA

Arnoldi-based orthonormal and hierarchical divergence-free polynomial basis and its applications

This paper presents a methodology to construct a divergence-free polynomial basis of an arbitrary degree in a simplex (triangles in 2D and tetrahedra in 3D) of arbitrary dimension. It allows for fast computation of all numerical solutions from degree zero to a specified degree \textit{k} for certain PDEs. The generated divergence-free basis is orthonormal, hierarchical, and robust in finite-precision arithmetic. At the core is an Arnoldi-based procedure. It constructs an orthonormal and hierarchical basis for multi-dimensional polynomials of degree less than or equal to \textit{k}. The divergence-free basis is generated by combining these polynomial basis functions. An efficient implementation of the hybridized BDM mixed method is developed using these basis functions. Hierarchy allows for incremental construction of the global matrix and the global vector for all degrees (zero to \textit{k}) using the local problem solution computed just for degree \textit{k}. Orthonormality and divergence-free properties simplify the local problem. PDEs considered are Helmholtz, Laplace, and Poisson problems in smooth domains and in a corner domain. These advantages extend to other PDEs such as incompressible Stokes, incompressible Navier-Stokes, and Maxwell equations.

math.NA

Response of a plate in turbulent channel flow: Analysis of fluid-solid coupling

The paper performs simulation of a rectangular plate excited by turbulent channel flow at friction Reynolds numbers of 180 and 400. The fluid-structure interaction is assumed to be one-way coupled, i.e, the fluid affects the solid and not vice versa. We solve the incompressible Navier Stokes equations using finite volume direct numerical simulation in the fluid domain. In the solid domain, we solve the dynamic linear elasticity equations using a time-domain finite element method. The obtained plate averaged displacement spectra collapse in the low frequency region in outer scaling. However, the high frequency spectral levels do not collapse in inner units. This spectral behavior is reasoned using theoretical arguments. We further study the sources of plate excitation using a novel formulation. This formulation expresses the average displacement spectra of the plate as an integrated contribution from the fluid sources within the channel. Analysis of the net displacement source reveals that at the plate natural frequencies, the contribution of the fluid sources to the plate excitation peaks in the buffer layer. The corresponding wall-normal width is found to be $\approx 0.75δ$. We analyze the decorrelated features of the sources using spectral Proper Orthogonal Decomposition (POD) of the net displacement source. We enforce the orthogonality of the modes in an inner product with a symmetric positive definite kernel. The dominant spectral POD mode contributes to the entire plate excitation. The contribution of the remaining modes from the different wall-normal regions undergo destructive interference resulting in zero net contribution. The envelope of the dominant mode further shows that the location and width of the contribution depend on inner and outer units, respectively.

physics.flu-dyn

Analysis of wall-pressure fluctuation sources from DNS of turbulent channel flow

The sources of wall-pressure fluctuations in turbulent channel flow are studied using a novel framework. The wall-pressure power spectral density (PSD) $(ϕ_{pp}(ω))$ is expressed as an integrated contribution from all wall-parallel plane pairs, $ϕ_{pp}(ω)=\int_{-δ}^{+δ}\int_{-δ}^{+δ}Γ(r,s,ω)\,\mathrm{dr}\,\mathrm{ds}$, using the Green's function. Here, $Γ(r,s,ω)$ is termed the net source cross spectral density (CSD) between two wall-parallel planes, $y=r$ and $y=s$ and $δ$ is the half channel height. Direct Numerical Simulation (DNS) data at friction Reynolds number of $180$ and $400$ are used to compute $Γ(r,s,ω)$. Analysis of the net source CSD, $Γ(r,s,ω)$ reveals that the location of dominant sources responsible for the premultiplied peak in the power spectra at $ω^+\approx 0.35$ and the wavenumber spectra at $λ^+\approx 200$ is in the buffer layer at $y^+\approx 16.5$ and $18.4$ for $Re_τ=180$ and $400$, respectively. The contribution from a wall-parallel plane (located at distance $y^+$ from the wall) to wall-pressure PSD is log-normal in $y^+$ for $ω^+>0.35$. A dominant inner-overlap region interaction of the sources is observed at low frequencies. Further, the decorrelated features of the wall-pressure fluctuation sources are analyzed using spectral Proper Orthogonal Decomposition (POD). We require the modes to be orthogonal in an inner product with a symmetric positive definite kernel. Spectral POD supports the case that the net source is composed of two decorrelated components - active (dominant mode) and inactive (remaining modes). The structure represented by the dominant POD mode at the premultiplied wall-pressure PSD peak inclines in the downstream direction. At the low-frequency linear PSD peak, the dominant mode resembles a large scale vertical pattern.

physics.flu-dyn

A variational level set methodology without reinitialization for the prediction of equilibrium interfaces over arbitrary solid surfaces

A robust numerical methodology to predict equilibrium interfaces over arbitrary solid surfaces is developed. The kernel of the proposed method is the distance regularized level set equations (DRLSE) with techniques to incorporate the no-penetration and mass-conservation constraints. In this framework, we avoid reinitialization typically used in traditional level set methods. This allows for a more efficient algorithm since only one advection equation is solved, and avoids numerical error associated with the re-distancing step. A novel surface tension distribution, based on harmonic mean, is prescribed such that the zero level set has the correct the liquid-solid surface tension value. This leads to a more accurate triple contact point location. The method uses second-order central difference schemes which facilitates easy parallel implementation, and is validated by comparing to traditional level set methods for canonical problems. The application of the method, in the context of Gibbs free energy minimization, to obtain liquid-air interfaces is validated against existing analytical solutions. The capability of our current methodology to predict equilibrium shapes over both structured and realistic rough surfaces is demonstrated.

physics.comp-ph

A parallel and streaming Dynamic Mode Decomposition algorithm with finite precision error analysis for large data

A novel technique based on the Full Orthogonalization Arnoldi (FOA) is proposed to perform Dynamic Mode Decomposition (DMD) for a sequence of snapshots. A modification to FOA is presented for situations where the matrix $A$ is unknown, but the set of vectors $\{A^{i-1}v_1\}_{i=1}^{N-1}$ are known. The modified FOA is the kernel for the proposed projected DMD algorithm termed, FOA based DMD. The proposed algorithm to compute DMD modes and eigenvalues i) does not require Singular Value Decomposition (SVD) for snapshot matrices $X$ with $κ_2(X) \ll 1/ε_m$, where $κ_2(X)$ is the 2-norm condition number of the snapshot matrix and $ε_m$ is the relative round-off error or machine epsilon, ii) has an optional rank truncation step motivated by round off error analysis for snapshot matrices $X$ with $κ_2(X) \approx 1/ε_m$, iii) requires only one snapshot at a time, thus making it a 'streaming' method even with the optional rank truncation step, iv) consumes less memory and requires less floating point operations to obtain the projected matrix than existing projected DMD methods and v) lends itself to easy parallelism as the main computational kernel involves only vector additions, dot products and matrix vector products. The new technique is therefore well-suited for DMD of large datasets on parallel computing platforms. We show both theoretically and using numerical examples that for FOA based DMD without rank truncation, the finite precision error in the computed projection of the linear mapping is $O(ε_mκ_2(X))$. The proposed method is also compared to existing projected DMD methods for computational cost, memory consumption and relative round off error. Error indicators are presented that are useful to decide when to stop acquiring new snapshots. The proposed method is applied to several examples of numerical simulations of fluid flow.

physics.comp-ph