Searcharxiv⌕ Search

arXiv subjects

Pratyai Mazumder

Publications and source records attributed to Pratyai Mazumder.

3 recordsLinked to original sources

FP64 Is All You Want, INT8 Is All You Need, FP4/6/8 Is All You Have

Ozaki scheme II emulates FP64 matrix products with INT8 ones through residues modulo pairwise coprime moduli, and variants for FP8 and FP4 have followed. We treat these schemes as one family and pose the choice of a scheme as a combinatorial program that minimizes the number of low-precision GEMMs. Given, for each modulus, a finite set of ways to compute products modulo it from low-precision GEMMs, we find the choice of moduli and ways with the fewest GEMMs, for any format, accumulator and inner dimension, and derive lower bounds on the GEMM count over the whole family. Applied to the formats of current GPUs, the method gives the first FP6 schemes, an FP8 scheme with fewer GEMMs than any previous one, and an FP4 scheme that the bounds show needs the fewest GEMMs of any scheme in the family whose moduli lie in a stated range. Implemented on three Blackwell GPUs, the INT8, FP8 and FP4 schemes run faster than native FP64, up to 83x on B300.

cs.DC↗

The Canonical Parallel Form as a Substrate for Parallelizing Compilers and Agentic Optimizers

Imperative code fixes an execution order the computation does not require, and a parallelizing compiler must prove which parts of that order it can remove. We introduce the Canonical Parallel Form (CPF), a device-neutral program state from which every ordering constraint our analyses prove unnecessary has been removed. CPF is reached by output-preserving normalization, by lifting semantic operations such as tensor contractions, and by deriving parallelism in three levels ordered by decidability: syntactic subscript tests, exact affine dependence tests over integer sets, and SMT queries over non-linear integer arithmetic, which also admit parallelism guarded behind a runtime check. Heuristics can then specialize the canonical form for each architecture. Across 248 loop-level reasoning kernels on an AMD MI300A, CPF reaches 4.4x over Numba on its 24 Zen 4 cores and 25.4x on its CDNA 3 GPU. Against the other auto-parallelizing optimizers, CPF is 2.9x faster on the CPU and 8.7x on the GPU than DaCe's own auto-parallelizer, 1.9x faster than Pluto, and 1.5x faster than PPCG on the affine subset of the kernels. Because the pipeline is deterministic, CPF also serves as an agent's starting source, cutting the token cost per kernel by up to a factor of 2.72x while leaving the achieved speed-up unchanged, since the coding agents reason less about parallelism.

cs.PF↗

Computing the Full Earth System at 1 km Resolution

We present the first-ever global simulation of the full Earth system at 1.25 km grid spacing, achieving highest time compression with an unseen number of degrees of freedom. Our model captures the flow of energy, water, and carbon through key components of the Earth system: atmosphere, ocean, and land. To achieve this landmark simulation, we harness the power of 8192 GPUs on Alps and 20480 GPUs on JUPITER, two of the world's largest GH200 superchip installations. We use both the Grace CPUs and Hopper GPUs by carefully balancing Earth's components in a heterogeneous setup and optimizing acceleration techniques available in ICON's codebase. We show how separation of concerns can reduce the code complexity by half while increasing performance and portability. Our achieved time compression of 145.7 simulated days per day enables long studies including full interactions in the Earth system and even outperforms earlier atmosphere-only simulations at a similar resolution.

physics.ao-ph↗