Searcharxiv⌕ Search

arXiv · 2610.08430

NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale

Abstract

Agentic reinforcement learning (RL) disaggregates training from rollout, so each policy update must reach the rollout clusters before the next batch. Transferring a full 1T checkpoint for such weight synchronization (refit) takes 87.5 min between two AWS regions. Measurements of BF16 training show that about 1% of weights change their stored values per step. Recent systems exploit this sparsity but fall short on placement, exactness, or efficiency: they reimplement placement rules, assemble full tensors, rebuild values arithmetically, or use a cross-cluster collective, and none fully recovers from mid-refit failures. We present NeMo-DCR (Delta-Compressed Refit), which sends only changes yet is bit-exact: receivers obtain the same parameter and buffer bits as a dense refit. For placement, fixed affine mappings project changes from training shards into the checkpoint's canonical coordinates, residual conversion covers the other changes, and the serving runtime's native loader places all changes in receiver storage. For exactness, compressible XOR masks carry affine changes whose projection and loader preserve stored bits, and overwrites carry the others. Receivers apply both in place, retries overwrite partial writes, and a joint commit binds the policy to the baseline for the next delta. For efficiency, object storage or a relay tree streams payloads during delta construction, without a cross-cluster collective. Even at 3% and 5% change rates, NeMo-DCR refits of 30B-1T models are 12-40$\times$ faster than a transport-only full-checkpoint reference. A 1T relay-tree refit at 3% takes 150 s instead of 87.5 min, making refits practical for cross-cluster agentic RL at trillion-parameter scale.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Songlin Jiang, Zhiyu Li, Terry Kong, Yu Yao, Youngeun Kwon, Bernard Nguyen, Ashwath Aithal, Mario Di Francesco. 2026-10-06. NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale. https://arxiv.org/abs/2610.08430

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

VeriNC: Finding Design Risks of In-Network Computing Systems

The emergence of programmable switches has brought in-network computing (INC) into the spotlight in recent years. By offloading computation directly onto the data transmission process, INC improves network utilization, reduces latency to sub-RTT levels, saves link bandwidth, and maintains throughput. However, INC disrupts the transparency of traditional networks, forcing developers to consider network exceptions like packet loss and out-of-order. If not properly handled, these exceptions can lead to violations of application properties, such as cache consistency and lock exclusion. Usual testing cannot exhaustively cover these exceptions, raising doubts about the correctness of INC systems and hindering their deployment in the industry. This paper presents VeriNC, the first general-purpose tool for verifying INC systems. VeriNC provides a high-level specification language and saves developers 67.2% lines of code on average. To help better understand the behavior of the system, VeriNC offers configurable network environments. VeriNC enables developers to express INC-specific correctness properties. VeriNC translates developer-specified systems into state transition representations, performs model checking to detect potential design risks, and reports violation traces to developers. We propose optimizations for INC-specific scenarios to address the challenge of state space explosion. We modeled INC systems across four application domains and identified design risks with VeriNC in seconds. VeriNC has also been adopted to guide the design of a new INC protocol. Based on our verification experience, we summarize lessons that help develop a correct INC protocol. We further reproduce them in real systems to confirm the validity of our verification result.

cs.DC↗

TorchGWAS 1.0: GPU-accelerated GWAS at scale

Imaging, molecular, and machine-learning workflows can generate thousands of quantitative phenotypes in a single cohort, creating substantial computational and output bottlenecks when testing traits individually. TorchGWAS is a GPU-accelerated framework that uses batched operations for high-throughput, covariate-adjusted linear association testing across large panels of quantitative phenotypes. Across 500,036 allele-harmonized tests, TorchGWAS t statistics agreed with PLINK 2.0. On an NVIDIA H100 80-GB GPU with a 48-core Intel Xeon Gold 6442Y host and measured disk read and write rates of 5.98 and 1.49 GB/s, respectively, median end-to-end times for 4.57 billion associations (8,931,083 variants by 512 phenotypes in 35,365 samples) were 28.46 s for BED, 29.13 s for hard-call PGEN, 51.48 s for BGEN, and 58.95 s for dosage PGEN, including writing 36.7 GB of binary summary statistics. TorchGWAS provides an efficient Python-based framework for parallel fixed-effect association screening at biobank scale.TorchGWAS is implemented in Python and distributed as a documented source repository at https://github.com/ZhiGroup/TorchGWAS.

cs.DC↗

AI-Assisted Computational Reproducibility on the FABRIC Testbed

Computational reproducibility remains difficult despite being central to scientific research. In this paper, we show how the international FABRIC testbed, combined with a large language model (LLM) coding agent through LoomAI, can simplify reproducing published experiments across multiple domains. We reproduced three case studies on FABRIC, covering BBR-family congestion-control evaluations, LAMMPS molecular dynamics scaling benchmarks on a CPU-only MPI cluster, and stress protein homeostasis genomics pipelines. Rather than focusing only on matching numerical outputs, we evaluate whether the reproduced experiments support the same scientific conclusions as the original studies. The AI assistant was effective in setting up the environment, adapting code, and debugging, but struggled with the analysis stages that lacked clearly defined workflows, which required human guidance to establish execution order and data dependencies. Across the case studies, the AI-assisted workflow reduced reproduction effort by roughly 4--6 times. We conclude with practical recommendations for improving AI-assisted reproducibility on research testbeds.

cs.DC↗