Searcharxiv⌕ Search

arXiv · 2609.31422

Towards Mitigating Fabricated Consensus: The Active Provenance Gate for Multi-Agent Debate Synthesis

Abstract

Large language model-based multi-agent debate (MAD) systems are being increasingly used as complex decision pipelines in distributed processes. However, their final synthesis phase still remains inadequately controlled. Even with detailed debate logs, summarizing models are prone to fabricating smoothly written debate consensus that is not grounded in the debate's history. To address this safety gap, this paper presents empirical research and studies if the introduction of active post-debate verification can mitigate the production of such factually unsupported summaries, while still providing valuable information. Furthermore, it is examined whether explicitly signalling divergence is preferable in the absence of a reliable compromise. The Active Provenance Gate (APG) is introduced as a post-debate verification layer that treats the source as a hard constraint, analysing the debate logs, auditing each claim, and applying self-correction. In crisis simulations, the self-healing mechanism more than doubles the average data Provenance Fidelity in difficult condition scenarios, before the strict gate blocks unsupported claims and generates divergence reports. In the human study, a vast majority of the users (over 75%) preferred a report explicitly stating failure in critical scenarios, despite most of them perceiving fabricated consensus from the baseline system as more fluent. Our main contribution is the transition of data origin tracing from passive logging to active conditional blocking before publication.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jakub Masłowski, Jarosław A. Chudziak. 2026-09-25. Towards Mitigating Fabricated Consensus: The Active Provenance Gate for Multi-Agent Debate Synthesis. https://arxiv.org/abs/2609.31422

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Towards Strategy-Level RSI for Skill-Augmented Agents: Learning When to Reuse Skills from Execution Feedback

Long-running agents accumulate reusable Skills, but a Skill that is semantically relevant to a task is not necessarily worth loading in the current state. We study the applicability question that arises once a candidate Skill is known: should it be loaded in the current state? We propose SkillApt, which uses matched WITH/WITHOUT Skill executions on the same task state as persistent evidence, estimates the conditional marginal utility of the Skill, and chooses LOAD or ABSTAIN accordingly. The base model, agent architecture, and Skill contents stay fixed; only the external deployment policy changes. We call this constrained setting strategy-level recursive self-improvement (Strategy-Level RSI). On 20 Skills and 160 held-out states, as paired evidence accumulates, SkillApt's task success rises from 81.9% under a cold start to 91.3%, matching a strong zero-shot LLM controller; yet SkillApt activates Skills on only 26.3% of states, versus 98.8% for the zero-shot controller. A hard-candidate study shows that non-optimal Skills mostly leave correctness unchanged while raising execution cost, and occasionally cause correctness harm. An ablation shows that a history recording only WITH success makes the policy load almost everywhere, whereas paired evidence substantially improves selectivity. These results indicate that relevance is not applicability: the main effect of execution evidence is not to make the model stronger but to change how existing Skills are deployed, moving the system from near-always loading to selective reuse.

cs.MA↗

Personalization Matters: Long-Horizon Conversation Agent with User-Centric Information in Online Shopping Interactions

Personalized conversational shopping requires maintaining preference consistency over multi-turn interactions, where users reveal constraints gradually. Existing approaches often rely on static profiles and do not explicitly control long-horizon interaction behavior. We propose a multi-agent, multimodal Retrieval-Augmented Generation (RAG) framework that decomposes dialogue state tracking, recommendation retrieval, preference-aware reasoning, and response generation, while integrating product metadata, product reviews, image-derived descriptions, and user historical reviews. To evaluate interaction-level quality, we adopt a trajectory-level protocol with four dimensions: Global Preference Consistency, Cumulative Information Synthesis, Interaction Trajectory, and Tone Consistency. On an Amazon Reviews 2023 benchmark, retrieval-enabled variants outperform a no-RAG baseline on automatic trajectory metrics (average 4.82 vs. 3.74). In a small real-user study ($n{=}5$), the Full variant achieves the highest mean overall rating (4.60 vs. 2.20 for Baseline), providing exploratory evidence that role decomposition plus user-centric retrieval improves perceived personalization.\footnote{Code and dataset are available at: https://github.com/RenaGao/Multimodel_RAG_Indexing

cs.MA↗

ConventionPlay: Capability-Limited Training for Robust Ad-Hoc Collaboration

Ad-hoc collaboration often requires agents to identify and adhere to some shared convention within a cooperative task. Existing work on reinforcement learning (RL) for ad-hoc collaboration focuses on training agents that adapt to the conventions established by their partners. These methods fail to consider the possibility that while some partners might follow only a single fixed convention, others may themselves be capable of adapting to multiple conventions. Here we present ConventionPlay, an RL-based approach that teaches agents to discover their partner's optimal convention by training against a learned population of partners that exhibit different degrees of adaptability across conventions. Some of these partners follow a single, fixed convention, while others are able to adapt to a subset of the possible conventions for the task in question. The existence of partners that support a limited subset of conventions forces agents trained against this population to actively probe their partner's capabilities, and steer their partner towards the most effective joint strategy that they are capable of following. Our experimental results demonstrate that agents trained via ConventionPlay achieve superior performance to existing ad-hoc collaboration methods against test populations of partners that are compatible with multiple conventions.

cs.MA↗