Searcharxiv⌕ Search

arXiv · 2610.09424

Democratizing MoE inference on commodity GPUs with CoMoE

Abstract

Deploying Mixture-of-Experts (MoE) models relies heavily on Expert Parallelism, which generates intense inter-GPU communication. Consequently, state-of-the-art inference systems require high-bandwidth, P2P interconnects (e.g., NVLink) in datacenter GPUs to handle massive token routing, making deployment prohibitively expensive. Consumer GPUs offer comparable compute power at significantly lower cost, promising to democratize MoE inference for individuals and enable privacy-preserving local deployments. However, their bandwidth-limited (only weak PCIe bus bandwidth) and host-mediated interconnects (no P2P support) introduce severe communication bottlenecks. We present CoMoE, a communication-efficient MoE inference system that resolves this mismatch through novel host-centric routing. Our key insight is that the unique communication topology provides the opportunity to elevate the host to an active routing hub, which can fundamentally reduce communication volume and eliminate global synchronization-induced stalls. Specifically, for token dispatch, we introduce host-backed token multicast to write shared tokens to the host exactly once, eliminating outbound transmission redundancy. For token combine, we propose a fine-grained, token-level aggregation mechanism using host staging buffers, which replaces rigid global synchronization and mitigates straggler effects. Evaluation on RTX 5090 GPUs shows that CoMoE improves inference throughput by up to 1.46x, approaching the performance of NVLink-capable A800 GPUs at only 23.4% of the hardware cost.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ruwen Fan, Yuezhi Zu, Junru Li, Qingda Hu, Xinjun, Yang, Jiwu Shu, Youyou Lu. 2026-10-07. Democratizing MoE inference on commodity GPUs with CoMoE. https://arxiv.org/abs/2610.09424

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

eAVID: Asynchronous Verifiable Information Dispersal with Post-Dissemination Pruning

The well-known Asynchronous Verifiable Information Dispersal (AVID) problem lets a sender disperse a message across $N=3F+1$ nodes such that it remains recoverable despite up to $F$ Byzantine failures and unbounded message delays. An optimal AVID scheme requires $3\times$ the message size in communication and storage: up to $F$ nodes may be delayed in responding, and among the remaining $2F+1$ respondents, up to $F$ may be Byzantine. This paper presents eAVID, an elastic AVID scheme that provides two key improvements. First, after dissemination, nodes can prune up to half of their stored information without requiring anyone to reconstruct or recode the message. Second, pruning is enabled by an extremely simple protocol. Nodes collect acknowledgments from one another confirming receipt of their coded pieces, after which each node locally prunes its stored information. eAVID achieves this $2\times$ reduction while making two tradeoffs: the sender generates $2N$ fragments rather than $N$, and pruning may necessitate contacting $2F+1$ nodes for message retrieval rather than $F+1$. eAVID was implemented in DispersedSimplex, a simple-to-understand BFT consensus protocol that erasure-codes its blocks. The implementation demonstrates two features. First, eAVID achieves these improvements using a flat erasure-coding scheme that requires no metadata or bookkeeping at the nodes during reconstruction. Second, it delivers the storage savings without a performance cost: with all nodes responsive, pruning fires on over $99\%$ of blocks and steady-state per-node storage falls by $42$-$46\%$ for committees of $10$ to $22$ nodes. Throughput and latency match the unmodified protocol showing that we can achieve these storage savings without any performance costs.

cs.DC↗

IAPRepair: In-Network Aggregation Enhanced Proactive Repair for Erasure-Coded Storage System

Erasure-coded storage provides fault tolerance with substantially lower storage overhead than full replication, but repairing lost or at-risk blocks requires intensive cross-node data transfer. Existing reactive repair schemes start only after a failure, while proactive schemes can move data before failure but commonly treat migration, reconstruction, and network aggregation as loosely coupled operations. As a result, receiver?side bottlenecks, heterogeneous available bandwidth, and limited programmable-switch state continue to constrain repair paral?lelism. This paper presents IAPRepair, an in-network aggre?gation enhanced proactive repair framework for erasure-coded storage. IAPRepair jointly constructs each repair batch, assigns reconstruction providers and replacement nodes according to normalized transmission loads, schedules migration around the remaining receive capacity, and selectively enables in-network aggregation for reconstruction blocks that would otherwise over?load healthy nodes. The selective design reduces receiver-side traffic while retaining migration parallelism and respecting a configurable switch-resource budget. We implement IAPRepair with a Tofino programmable switch and 16 storage nodes, and evaluate it using both a prototype testbed and large-scale simu?lations. Across coding parameters, block sizes, node populations, and multiple STF-node scenarios, IAPRepair reduces repair time by at least 47.83% compared with the evaluated state-of-the-art methods. The results demonstrate that coordinating proactive repair decisions with in-network processing is an effective way to improve repair efficiency under bandwidth heterogeneity.

cs.DC↗

Zepp: Accelerating Distributed MoE Serving under Relaxed Balance Constraints

As Mixture-of-Experts (MoE) models continue to scale, serving them increasingly relies on expert parallelism (EP) across a growing number of devices. Yet skewed expert workloads create imbalance across computation, communication, and memory, making load balancing a central optimization objective in distributed MoE serving. We observe that balance is not free: operations introduced to balance one dimension can themselves be expensive or imbalanced. This motivates us to rethink balance as a constraint rather than an optimization objective. We present Zepp, which directly optimizes the bottleneck communication in distributed MoE serving subject to simplified balance constraints on physical resources, i.e., GPUs and NICs. Zepp progressively optimizes inter-node communication across placement, routing, and execution. It first places expert replicas to reduce token communication under GPU constraints, then reshapes communication flows through split and merge primitives under NIC constraints, and finally partitions and schedules ex- pert computation to overlap the resulting communication. To adapt to dynamic workloads, Zepp jointly coordinates computation, token communication, and expert-weight movement at each iteration. Together, these designs allow Zepp to pursue the most efficient execution rather than a single-dimension balanced one. We implement Zepp and evaluate it against 7 state-of-the-art MoE serving systems, achieving up to 6.68$\times$ MoE layer speedup and a geometric mean speedup of 1.86$\times$ over the fastest competing baseline.

cs.DC↗