Searcharxiv⌕ Search

arXiv · 2610.03203

AFORE: Attention-FFN Disaggregation with Overlapped Reconfiguration of Experts

Abstract

Efficient serving of Mixture-of-Experts (MoE) models is challenging due to large expert parameters, input-dependent expert activation, and dynamic workloads. Expert parallelism distributes expert computation across GPUs, while attention-FFN disaggregation (AFD) separates attention and feed-forward computation into independent worker pools. However, we observe that a naive AFD implementation could make expert load imbalance more harmful: once FFN computation becomes an independent pipeline stage, overloaded experts directly slow the FFN stage and degrade end-to-end serving performance. To solve this problem, we present AFORE, a timely expert reconfiguration system for AFD-based MoE serving. AFORE exploits two architectural properties of AFD. First, AFD exposes the expert-token distribution of upcoming microbatches before they reach FFN execution, enabling placement decisions based on near-future demand instead of stale historical profiles. Second, AFD creates a pipeline window in which expert migration for a target microbatch can be overlapped with the computation of preceding in-flight microbatches. AFORE formulates expert reconfiguration as a microbatch-aware scheduling problem and uses a migration-aware scheduler to decide when and which experts to migrate. AFORE further implements lightweight demand prefetching and NVLink-based GPU-GPU expert migration to reduce reconfiguration overhead. Evaluation on a 110B-parameter MoE model across four dynamic workloads shows that AFORE improves output throughput by 10.1-17.6% and reduces P95 inter-token latency by 7.1-9.5% compared with the strongest competing baseline. Compared with static placement, AFORE improves throughput by 29.8% on average and reduces P95 inter-token latency by 18.2% on average. Migration profiling further shows that AFD pipeline overlap can fully hide expert-migration latency.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wenshuang Li, Youhe Jiang, You Peng, Jiawei Jiang, Binhang Yuan. 2026-10-02. AFORE: Attention-FFN Disaggregation with Overlapped Reconfiguration of Experts. https://arxiv.org/abs/2610.03203

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Grassroots Bonds: Financing by the People, for the People

Grassroots currencies turn mutual trust into liquidity: a grassroots coin is a unit of its issuer's debt, backed by the issuer's goods and services, which the issuer must redeem, 1-for-1, against any coin they hold; liquidity arises from mutual credit lines, formed by the voluntary exchange of coins among persons who trust each other. As coins are redeemable 1-for-1, the exchange must be 1-for-1 as well, lest prompt redemption after it leave one party with undue profit. Thus, grassroots coins are incongruent with interest-bearing credit. Here, we extend grassroots currencies to include also grassroots bonds, units of their issuer's debt due at a later date. Upon maturity, the bearer of the bond may redeem it against a coin of the issuer, so liquid coins can be lent against interest-bearing bonds. We specify bonds by adding Date and Escrow clauses to the social contract of grassroots currencies: the contract enforces the redemption of a bond when it is mature according to the date stated by the issuer, who undertakes to keep it current. We show that the voluntary swap of coins and bonds can realise the basic financial instruments: loans, sale of debt, and forward contracts, and, with an escrow agent, loans with payment schedules, options, collateral, guarantees, insurance, credit default swaps, letters of credit, and credit lines. We extend the liquidity ratios of corporate finance to bonds, with maturity as asset class, and prove that a community clears its debts by redemptions exactly when no member owes more than they hold, and that it does so without coordination once dates advance and persons act on their rights. Grassroots currencies that include coins and bonds are implemented in GLP, a concurrent logic programming language running on Dart, as a program derived from the contract and demonstrated by a village market of six agents and an escrow agent.

cs.DC↗

Characterization-Guided GPU Fault Resilience in NVIDIA MPS

NVIDIA Multi-Process Service (MPS) enables fine-grained GPU sharing by allowing multiple processes to execute concurrently on the same GPU, making it an important mechanism for improving GPU utilization. However, MPS has weak fault resilience: a fault in one process can terminate all co-running processes, limiting its adoption in resilience-critical settings such as multi-tenant GPU clusters. In this work, we design fault-resilient MPS to solve this problem. Our design is guided by insights from a systematic characterization of GPU faults and a deep analysis of their end-to-end processing pipeline. Based on these insights, we design two complementary mechanisms. First, we design a fault isolation mechanism for the dominant memory-related faults that can be fully isolated while preserving process-level fail-stop semantics by software intervention in the open GPU driver kernel module. For other faults whose process is within proprietary software, we design a fast-recovery substrate that combines virtual-memory-based GPU-resident state sharing with pre-initialized standbys. Our evaluation across GPUs and workloads demonstrates effective fault isolation and fast recovery with minimal overhead: in an end-to-end case study, isolation incurs no visible outage, while recovery restores pre-fault throughput in 355\,ms.

cs.DC↗

ARGOS: Reinforcement Learning-Driven Multidimensional Elasticity for Service Orchestration in the Computing Continuum

Data-intensive services in the Computing Continuum must balance analytics quality, resource usage, and cost across heterogeneous nodes with limited and uneven capacity. This balance becomes especially difficult when resource scaling reaches capacity limits, because changes in demand and cluster pressure must then be absorbed without violating client-defined quality ranges. Existing orchestrators mainly adapt resources, placements, or replicas, while analytics requirements such as coverage, sample, and freshness remain fixed. This article presents ARGOS, the Adaptive Reinforcement Learning-Driven Governance for Orchestrated Services, an end-to-end controller that formulates multidimensional elasticity as a per-request Markov decision process over analytics quality and cluster pressure, supported by capacity-aware admission. ARGOS is evaluated under controlled workloads and time-varying multi-tenant arrivals on a heterogeneous cluster. Across the controlled scenarios, the deep reinforcement learning policies consistently outperform the non-learning baselines and approach the independently tuned best-fixed reference. A separate live evaluation reports improvements over the static midpoint under realistic and saturated arrivals, with no recorded CPU or memory violations but remaining coverage violations. These results support deep reinforcement learning as an adaptive mechanism for multidimensional elasticity when resource scaling alone is insufficient.

cs.DC↗