Searcharxiv⌕ Search

arXiv · 2610.10100

The Cost of Classical Multi-Agent Path Finding

Abstract

Multi-Agent Path Finding (MAPF) is the problem of planning conflict-free paths for multiple agents in a shared space, each from its start to its goal. Classical MAPF has been the dominant formulation for many years, with its assumptions of discrete time and graph-based conflicts presumably easing the search for solutions. These assumptions limit the physical environments and agents for which a solution is truly collision-free, and also place an upper bound on solution quality that no algorithmic improvements can lift. This work investigates how much solution quality, and in what contexts, the classical MAPF formulation forfeits. Continuous-time MAPF (MAPF$_R$) relaxes these assumptions, making it a natural counter-formulation to compare against across various agent counts and sizes, and graph connectedness, topologies, and resolutions. We find that continuous time and agent shape consideration are worth relatively little on their own; their value comes from enabling an expanded range of move actions, on average improving solution quality by at least $5\%$ on narrow and constrained maps and $17\%$ on maps with open spaces. In some cases, the improvements exceed $20\%$. Doubling the map resolution with classical MAPF recovers less than $3\%$, meaning that little of what is forfeited can be bought back through more compute. This work therefore provides insight on when classical MAPF is a reasonable simplification, and when MAPF$_R$ unlocks significantly higher-quality solutions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Alvin Combrink, Sabino Francesco Roselli, Martin Fabian. 2026-10-07. The Cost of Classical Multi-Agent Path Finding. https://arxiv.org/abs/2610.10100

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Towards Strategy-Level RSI for Skill-Augmented Agents: Learning When to Reuse Skills from Execution Feedback

Long-running agents accumulate reusable Skills, but a Skill that is semantically relevant to a task is not necessarily worth loading in the current state. We study the applicability question that arises once a candidate Skill is known: should it be loaded in the current state? We propose SkillApt, which uses matched WITH/WITHOUT Skill executions on the same task state as persistent evidence, estimates the conditional marginal utility of the Skill, and chooses LOAD or ABSTAIN accordingly. The base model, agent architecture, and Skill contents stay fixed; only the external deployment policy changes. We call this constrained setting strategy-level recursive self-improvement (Strategy-Level RSI). On 20 Skills and 160 held-out states, as paired evidence accumulates, SkillApt's task success rises from 81.9% under a cold start to 91.3%, matching a strong zero-shot LLM controller; yet SkillApt activates Skills on only 26.3% of states, versus 98.8% for the zero-shot controller. A hard-candidate study shows that non-optimal Skills mostly leave correctness unchanged while raising execution cost, and occasionally cause correctness harm. An ablation shows that a history recording only WITH success makes the policy load almost everywhere, whereas paired evidence substantially improves selectivity. These results indicate that relevance is not applicability: the main effect of execution evidence is not to make the model stronger but to change how existing Skills are deployed, moving the system from near-always loading to selective reuse.

cs.MA↗

Personalization Matters: Long-Horizon Conversation Agent with User-Centric Information in Online Shopping Interactions

Personalized conversational shopping requires maintaining preference consistency over multi-turn interactions, where users reveal constraints gradually. Existing approaches often rely on static profiles and do not explicitly control long-horizon interaction behavior. We propose a multi-agent, multimodal Retrieval-Augmented Generation (RAG) framework that decomposes dialogue state tracking, recommendation retrieval, preference-aware reasoning, and response generation, while integrating product metadata, product reviews, image-derived descriptions, and user historical reviews. To evaluate interaction-level quality, we adopt a trajectory-level protocol with four dimensions: Global Preference Consistency, Cumulative Information Synthesis, Interaction Trajectory, and Tone Consistency. On an Amazon Reviews 2023 benchmark, retrieval-enabled variants outperform a no-RAG baseline on automatic trajectory metrics (average 4.82 vs. 3.74). In a small real-user study ($n{=}5$), the Full variant achieves the highest mean overall rating (4.60 vs. 2.20 for Baseline), providing exploratory evidence that role decomposition plus user-centric retrieval improves perceived personalization.\footnote{Code and dataset are available at: https://github.com/RenaGao/Multimodel_RAG_Indexing

cs.MA↗

ConventionPlay: Capability-Limited Training for Robust Ad-Hoc Collaboration

Ad-hoc collaboration often requires agents to identify and adhere to some shared convention within a cooperative task. Existing work on reinforcement learning (RL) for ad-hoc collaboration focuses on training agents that adapt to the conventions established by their partners. These methods fail to consider the possibility that while some partners might follow only a single fixed convention, others may themselves be capable of adapting to multiple conventions. Here we present ConventionPlay, an RL-based approach that teaches agents to discover their partner's optimal convention by training against a learned population of partners that exhibit different degrees of adaptability across conventions. Some of these partners follow a single, fixed convention, while others are able to adapt to a subset of the possible conventions for the task in question. The existence of partners that support a limited subset of conventions forces agents trained against this population to actively probe their partner's capabilities, and steer their partner towards the most effective joint strategy that they are capable of following. Our experimental results demonstrate that agents trained via ConventionPlay achieve superior performance to existing ad-hoc collaboration methods against test populations of partners that are compatible with multiple conventions.

cs.MA↗