Searcharxiv⌕ Search

arXiv · 2609.40069

Conversational Capture: A Trajectory-Level Framework for Evaluating Generative Engine Optimization in Multi-turn Human-Agent Interaction

Abstract

Generative Engine Optimization (GEO) shapes content to increase its likelihood of being cited by answer engines built on retrieval-augmented large language models. GEO is typically evaluated as a single-turn property: for a fixed query, an evaluator measures a source's visibility in one answer. We argue that the single answer is an inadequate unit of analysis. Human-agent information seeking forms a closed loop: the agent's answer changes the user's beliefs and therefore the next question, which in turn determines what the agent retrieves. We introduce conversational capture, a phenomenon in which a source cited early becomes substantially more likely to be cited again. Capture operates through a machine-side channel, history-conditioned retrieval, and a human-side channel, follow-up questions directed toward the captured source. We formalize the interaction as a two-layer closed-loop system and derive trajectory-level constructs: cumulative conversational visibility; a direct/feedback decomposition of trajectory gain; a nested split of the feedback term into machine-side and human-side channels; a capture coefficient; a compounding ratio; and a misranking diagnostic. Using reinforcement-process (Pólya-urn) theory, we prove that the feedback term is zero under single-turn evaluation and that GEO's cumulative payoff grows superlinearly with conversation length while capture develops. A model-derived illustration shows that the feedback term can exceed the direct term, the compounding ratio exceeds two within ten turns, and single-turn and trajectory rankings agree only weakly (Kendall's $τ= 0.4$). We connect the human channel to information foraging, trust calibration, and Bayesian persuasion, and discuss design implications for answer engines.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Junwei Yu, Jieyu Zhou, Mufeng Yang, Yepeng Ding, Hiroyuki Sato. 2026-09-30. Conversational Capture: A Trajectory-Level Framework for Evaluating Generative Engine Optimization in Multi-turn Human-Agent Interaction. https://arxiv.org/abs/2609.40069

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

How University Disability Services Professionals Write Image Descriptions for HCI Figures Using Generative AI

Disability Services Office (DSO) professionals at higher education institutions write alt text for visual content. However, due to the complexity of visual content, such as HCI figures in research publications, DSO professionals can struggle to write high-quality alt text if they lack subject expertise. Generative AI has shown potential for understanding figures and writing descriptions, yet its support for DSO professionals remains underexplored, and few studies evaluate the quality of AI-assisted alt text. In this work, we conducted two studies: first, we investigated generative AI support for writing alt text for HCI figures with 12 DSO professionals. Second, we recruited 11 HCI experts to evaluate the alt text written by DSO professionals. Findings show that alt text written solely by DSO professionals has lower quality than alt text written with AI assistance. AI assistance also helped DSO professionals write alt text more quickly and with greater confidence; however, they reported inefficacy in interactions with the AI. Our work contributes to research on AI support for non-subject-expert accessibility professionals.

cs.HC↗

"AI Psychosis" in Context: How Conversation History Shapes LLM Responses to Delusional Beliefs

Extended interaction with large language models (LLMs) has been linked to the reinforcement of delusional beliefs, attracting clinical and public concern. Yet most empirical work evaluates model safety in brief interactions, which may not reflect how harms develop through sustained dialogue. Five LLMs were tested across three levels of accumulated context, using the same escalating delusional conversation history to isolate its effect on model behaviour. Responses were coded on risk and safety dimensions, and each model was analysed qualitatively. Models separated into two distinct tiers: GPT-4o, Grok 4.1 Fast, and Gemini 3 Pro exhibited high-risk, low-safety profiles; Claude Opus 4.5 and GPT-5.2 Instant displayed the opposite pattern. As context accumulated, performance degraded in the unsafe group, while the same material activated stronger safety interventions among safer models. Qualitative analysis identified distinct mechanisms of failure, including validating the user's delusional premises, elaborating beyond them with new content, and attempting harm reduction from within the delusional frame. Safer models, however, often used the established relationship to support intervention, challenging delusional beliefs and directing the user to external support. These findings indicate that accumulated context functions as a stress test of safety architecture, revealing whether prior dialogue is treated as a worldview to inherit or evidence to evaluate. Short-context assessments may therefore mischaracterise model safety, underestimating danger in some systems while missing context-activated gains in others. The results suggest that delusion reinforcement is a tractable alignment failure, with safer models establishing a baseline that future systems should now be expected to meet.

cs.HC↗

"ChatGPT, help me draft a breakup text": The Covert Triad and Articulation Labor in AI-Assisted Romantic Communication

Generative artificial intelligence (AI) has begun infiltrating the most ordinary domains of romantic life---drafting apologies, softening reproaches, and decoding a partner's ambiguous messages. While recent scholarship on AI in intimate life has concentrated on chatbot companions, this article shifts the frame to AI as an intermediary in human-to-human romantic communication. Drawing on a multimodal corpus of public accounts, commentary, and representations from 2023 to 2026, we contribute two complementary concepts. The covert triad names a structural reconfiguration: a relationship that appears dyadic but operates triadically, with AI visible only to the partner who deploys it. Articulation labor identifies the mechanism whereby the expressive component of emotional labor---converting felt experience into language a partner can receive---is delegated to AI, even as feeling labor remains lodged in the user. Authenticity thus moves from linguistic authorship toward emotional ownership, a change actively contested within a distinctly therapeutic model of intimacy.

cs.HC↗