Searcharxiv⌕ Search

arXiv · 2610.09657

Fast and Memory Efficient Offload Training Framework with Hybrid XPU Computation

Abstract

With the ever-growing size of deep learning models, GPU memory is prone to being insufficient during training. A prominent approach is ZeRO-Offload, which moves the optimizer states to CPU memory and performs parameter update using CPU. However, the deficiencies of ZeRO-Offload include low GPU utilization, imperfect overlapping of communication and computation, and inflexible offloading. In this paper, we leverage Direct Host Access (DHA) on the GPU that can compute data in CPU memory, forming a novel hybrid on-GPU and DHA. We design and implement MemFerry consisting of an execution scheduler and a shadow model. The scheduler strategically chooses layers of parameters for DHA computation and transmits the remaining parameters to GPU memory simultaneously to shorten forward propagation time, and further loads DHA parameters to GPU memory to reduce backward propagation time. The shadow model presents a unified memory abstraction for the parameter partitions stored separately in GPU and CPU memories. To further reduce GPU memory usage, we present MemFerry along with its dynamic programming algorithm that offloads gradients to CPU memory via DHA. We further extend MemFerry to emerging scale-up domains with ScaleUp-MemFerry, which exploits otherwise underutilized accelerator interconnect bandwidth to assist host-to-accelerator data movement through adaptive multi-path transfer. Our experiments show that \system trains up to $1.68\times$ faster and MemFerry can train $1.52\times$ larger model compared to ZeRO-Offload on a single GPU, and increase training speed by at least $28.1\%$ when scaling to data parallelism on 8 GPUs. We further extend the design to a Huawei CloudMatrix384 scale-Up node with up to 8 NPUs, and our ScaleUp-MemFerry reduces the end-to-end iteration time by up to $20.7\%$ over DeepSpeed.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhiyi Yao, Zuning Liang, Yuedong Xu, Jin Zhao, Jessie Hui Wang, Tong Li. 2026-10-07. Fast and Memory Efficient Offload Training Framework with Hybrid XPU Computation. https://arxiv.org/abs/2610.09657

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

eAVID: Asynchronous Verifiable Information Dispersal with Post-Dissemination Pruning

The well-known Asynchronous Verifiable Information Dispersal (AVID) problem lets a sender disperse a message across $N=3F+1$ nodes such that it remains recoverable despite up to $F$ Byzantine failures and unbounded message delays. An optimal AVID scheme requires $3\times$ the message size in communication and storage: up to $F$ nodes may be delayed in responding, and among the remaining $2F+1$ respondents, up to $F$ may be Byzantine. This paper presents eAVID, an elastic AVID scheme that provides two key improvements. First, after dissemination, nodes can prune up to half of their stored information without requiring anyone to reconstruct or recode the message. Second, pruning is enabled by an extremely simple protocol. Nodes collect acknowledgments from one another confirming receipt of their coded pieces, after which each node locally prunes its stored information. eAVID achieves this $2\times$ reduction while making two tradeoffs: the sender generates $2N$ fragments rather than $N$, and pruning may necessitate contacting $2F+1$ nodes for message retrieval rather than $F+1$. eAVID was implemented in DispersedSimplex, a simple-to-understand BFT consensus protocol that erasure-codes its blocks. The implementation demonstrates two features. First, eAVID achieves these improvements using a flat erasure-coding scheme that requires no metadata or bookkeeping at the nodes during reconstruction. Second, it delivers the storage savings without a performance cost: with all nodes responsive, pruning fires on over $99\%$ of blocks and steady-state per-node storage falls by $42$-$46\%$ for committees of $10$ to $22$ nodes. Throughput and latency match the unmodified protocol showing that we can achieve these storage savings without any performance costs.

cs.DC↗

IAPRepair: In-Network Aggregation Enhanced Proactive Repair for Erasure-Coded Storage System

Erasure-coded storage provides fault tolerance with substantially lower storage overhead than full replication, but repairing lost or at-risk blocks requires intensive cross-node data transfer. Existing reactive repair schemes start only after a failure, while proactive schemes can move data before failure but commonly treat migration, reconstruction, and network aggregation as loosely coupled operations. As a result, receiver?side bottlenecks, heterogeneous available bandwidth, and limited programmable-switch state continue to constrain repair paral?lelism. This paper presents IAPRepair, an in-network aggre?gation enhanced proactive repair framework for erasure-coded storage. IAPRepair jointly constructs each repair batch, assigns reconstruction providers and replacement nodes according to normalized transmission loads, schedules migration around the remaining receive capacity, and selectively enables in-network aggregation for reconstruction blocks that would otherwise over?load healthy nodes. The selective design reduces receiver-side traffic while retaining migration parallelism and respecting a configurable switch-resource budget. We implement IAPRepair with a Tofino programmable switch and 16 storage nodes, and evaluate it using both a prototype testbed and large-scale simu?lations. Across coding parameters, block sizes, node populations, and multiple STF-node scenarios, IAPRepair reduces repair time by at least 47.83% compared with the evaluated state-of-the-art methods. The results demonstrate that coordinating proactive repair decisions with in-network processing is an effective way to improve repair efficiency under bandwidth heterogeneity.

cs.DC↗

Zepp: Accelerating Distributed MoE Serving under Relaxed Balance Constraints

As Mixture-of-Experts (MoE) models continue to scale, serving them increasingly relies on expert parallelism (EP) across a growing number of devices. Yet skewed expert workloads create imbalance across computation, communication, and memory, making load balancing a central optimization objective in distributed MoE serving. We observe that balance is not free: operations introduced to balance one dimension can themselves be expensive or imbalanced. This motivates us to rethink balance as a constraint rather than an optimization objective. We present Zepp, which directly optimizes the bottleneck communication in distributed MoE serving subject to simplified balance constraints on physical resources, i.e., GPUs and NICs. Zepp progressively optimizes inter-node communication across placement, routing, and execution. It first places expert replicas to reduce token communication under GPU constraints, then reshapes communication flows through split and merge primitives under NIC constraints, and finally partitions and schedules ex- pert computation to overlap the resulting communication. To adapt to dynamic workloads, Zepp jointly coordinates computation, token communication, and expert-weight movement at each iteration. Together, these designs allow Zepp to pursue the most efficient execution rather than a single-dimension balanced one. We implement Zepp and evaluate it against 7 state-of-the-art MoE serving systems, achieving up to 6.68$\times$ MoE layer speedup and a geometric mean speedup of 1.86$\times$ over the fastest competing baseline.

cs.DC↗