Searcharxiv⌕ Search

arXiv · 2609.26080

The Fleet Is the Model: Engineering Collective Intelligence with Fusion-MoA Pioneer R1

Abstract

The model exposed to an application need not be a single checkpoint; it can be a governed fleet. Existing serving systems manage checkpoints and replicas, while multi-agent frameworks compose model calls without defining a stable collective identity, effect authority, or member-level evolution. We present Fusion-MoA, a runtime that exposes independently served heterogeneous Cells as one OpenAI-compatible model. A versioned Profile controls membership and evidence admission; read-only Analysts contribute bounded evidence, while a sole Executor retains all final-answer and tool authority. Cells can be qualified, promoted, or rolled back without changing the public API. We evaluate an eight-Cell, three-lineage deployment through three operational witnesses. On a fixed HMMT P1-P10 slice, the collective solves 8/10 problems versus 6/10 for the strongest individual Cell, and a preserved trace shows minority knowledge transferred to three initially incorrect or empty Cells. On 20 Terminal-Bench 2.1 tasks, all tool actions remain attributable to one Executor, with zero Analyst actions and zero bypass effects. Six Cells are promoted and one incompatible candidate is locally rolled back while the service remains available. Fusion-MoA demonstrates that heterogeneous model capability can be operated as one observable, authority-bounded, and independently evolvable service.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zongyou Yang, Yinghan Hou. 2026-08-03. The Fleet Is the Model: Engineering Collective Intelligence with Fusion-MoA Pioneer R1. https://arxiv.org/abs/2609.26080

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

REAT: A Reflective Experience-Augmented Tutoring Framework for Multi-turn Mathematical Instruction

Current Large Language Models (LLMs) excel at solving complex mathematical problems, yet this proficiency does not inherently translate into effective tutoring. While advanced LLM tutors may leverage multi-agent frameworks or fine-tuning, most still lack a mechanism to systematically accumulate and reuse pedagogical experience over time, limiting their adaptability to diverse student needs during fluid, multi-turn interactions. To bridge this gap, we propose the Reflective Experience-Augmented Tutoring (REAT) framework, which couples experience distillation from historical dialogues with real-time adaptive retrieval. Driven by a multi-agent Observer-Critic-Mentor (OCM) distillation pipeline, REAT reviews past conversational trajectories and distills raw interactions into structured, problem-agnostic pedagogical experiences. During live tutoring, a state-aware retrieval module injects these curated experiences to provide adaptive scaffolding based on the student's cognitive state. Experiments demonstrate that the proposed framework significantly outperforms both prompt-only and supervised fine-tuning (SFT) baselines, particularly in improving complex, low-scoring tutoring scenarios. Crucially, the distilled experiences exhibit robust generalization across diverse model architectures and mathematical datasets.

cs.MA↗

Optimization without Future Compromises? Decentralized Coordination via Collective and Reinforcement Learning

Efficient resource allocation in multi-agent systems requires autonomous agents to coordinate their decisions while balancing system-wide objectives with individual costs. This becomes increasingly challenging over long time horizons, where decisions that improve the current allocation may compromise future resource allocation, while decentralized agents have limited observations of the overall system. Multi-agent reinforcement learning (MARL) can learn such long-term dependencies via local observations, but directly applying it to large-scale coordination leads to rapidly growing decision spaces and inefficient training. To this end, we propose Hierarchical Reinforcement and Collective Learning (HRCL), a hierarchical framework that uses MARL to guide, rather than replace, decentralized multi-agent coordination. At the high level, MARL learns strategies that restrict the alternatives considered during coordination and guide agents in balancing system-wide and individual objectives. At the low level, agents perform efficient decentralized coordination under this strategic guidance. This separation reduces the learning space and allows short-term coordination trade-offs to be evaluated according to their long-term effects. Experiments on a synthetic benchmark show that HRCL converges substantially faster than standalone MARL and reduces system-wide and individual costs by 35.53% and 27.05%, respectively. Evaluations on energy self-management and drone swarm sensing further show improved resource allocation, power-peak regulation, and sensing efficiency. These results show that learning strategic guidance for an existing coordination process can retain scalable decentralized coordination without letting short-term decisions compromise future resource allocation.

cs.MA↗

Multi-robot Graph Traversal with Support Coordination under Stochastically Moving Adversaries

Cooperative multi-robot missions require team of robots to traverse environments where adversaries or hazards with stochastic dynamics induce time-varying traversal risk. While support coordination--where robots assist teammates in traversing risky regions--can significantly reduce mission costs, its effectiveness depends on the team's ability to anticipate future risk. We formulate support-based multi-robot graph traversal problem with stochastically moving adversaries, where future risky regions become uncertain as adversaries move through the environment. When adversaries remain stationary, our formulation reduces to the static risky-edge setting. To address the stochastic case, we model individual adversaries as first-order Markov stay-move processes over graph edges and propagate their occupancy distributions over a finite planning horizon to obtain time-indexed edge-risk forecasts. These forecasts inform the support candidate selection and joint robot path planning. Experimental results show that forecast-informed support decisions consistently lower expected team cost relative to evaluated baselines in stochastic motion settings.

cs.MA↗