Searcharxiv⌕ Search

arXiv · 2610.07732

Who Is Talking to the Agent? LLMs in Multi-User 3D Virtual Environments

Abstract

When several people share a 3D virtual room with an LLM agent, the agent must decide not only what to say, but whether an utterance was addressed to it and, if accessible, what profile information about the others present it may use. To study both problems, we construct LookAway, a controlled corpus of 40 sessions involving 80 distinct personas and an LLM agent (1,200 turns), including ambiguous-addressee turns in which speaker orientation agrees or conflicts with the intended addressee. Across three open-weight large language models and five conditions varying which profiles the agent sees and whether it is told where each person faces (18,000 decisions), adding speaker orientation increased addressee accuracy from 56% to 99.5% when orientation was congruent, but when the speaker faced someone other than the addressee, two of the models went by where the speaker faced on more than 85% of those turns. Warning one model that orientation could be misleading reduced this only modestly. A browser-based 3D demonstrator shows the effect live: the same sentence gets an answer when the speaker faces the agent and silence when they face the other person. Providing both personas' profiles improved responses about the person being asked about, but also increased the use of profile attributes not revealed in the shared conversation, reaching 45.3% of answers for one model. More context thus improves multi-user interaction but also leads to oversharing, so shared LLM agents need mechanisms for weighing spatial cues and controlling when user-specific information enters a response.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mohammad Al-Ratrout, Shayla Sharmin, Roghayeh Leila Barmaki. 2026-10-06. Who Is Talking to the Agent? LLMs in Multi-User 3D Virtual Environments. https://arxiv.org/abs/2610.07732

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

CHOMP: Multimodal Chewing Side Detection with Earphones

Chewing-side preference (CSP) is a risk factor for temporomandibular disorders (TMDs) and a behavioral manifestation. Although TMDs affect roughly one-third of the global population, assessment relies on clinical examinations and self-reports, providing limited insight into everyday jaw function. We present CHOMP, the first earphone-based chewing-side detection system for continuous CSP monitoring. Using OpenEarable 2.0, we collected multimodal data from 20 participants with microphones, a bone-conduction microphone, IMU, PPG, and a pressure sensor across diverse foods, activities, and acoustic-interference conditions. CHOMP models paired-ear temporal feature sequences using modality-specific bidirectional GRUs, multimodal fusion, and prototype-based classification with an optional short user adaptation. Microphones achieve the strongest single-sensor performance, with median macro F1 scores of 97.2% under leave-one-food-out (LOFO) and 95.7% under leave-one-subject-out (LOSO) evaluation after user adaptation. Multimodal fusion reaches 98.0% under LOFO and 97.3% under adapted LOSO. We demonstrate CHOMP's performance under three acoustic-interference conditions and within a cafeteria setting. Our results establish earphones as a practical platform for everyday CSP monitoring and jaw-function assessment.

cs.HC↗

How Children Design and Reason about Trustworthy AI Chatbots

Children increasingly interact with AI chatbots, making trust calibration essential to AI literacy. Prior research has examined children's trust in AI mainly as users evaluating systems built by others, rather than as designers of their own chatbots. We developed a chatbot-building environment with adjustable trust-relevant traits (e.g., confidence, transparency, formality, assertiveness), rules, and persona. We conducted mixed-methods study with 115 learners (ages 8-18) who made 119 chatbots. We examined how children configured their chatbots, reasoned about trustworthiness, and how closely chatbot behavior aligned with their designs. Younger students (age 10-13) set significantly higher confidence than older students (age 14-18), and some deliberately built chatbots that gave wrong answers on purpose, yet still called them trustworthy, arguing that a chatbot does what it was built to do. Younger students equated trust with purpose-fulfillment, while older students linked it to transparent, calibrated design. Students also calibrated academic chatbots to be more transparent and formal than hobby chatbots. We identify seven design dimensions describing what children believe makes a chatbot trustworthy, and discuss implications for AI literacy tools.

cs.HC↗

LeanSide: A Formally Verified Co-Reasoning System for Natural-language Proofs

Large language models are increasingly used as collaborators on deductive-reasoning tasks, but their outputs can hallucinate or pull users away from intended reasoning. Formal proof assistants provide machine-checked verification, but have a steep learning curve and require more granular reasoning than human written proofs. We explore an interface that combines these strengths, allowing users to write and revise free-form natural-language proofs while a verified backend checks their reasoning and returns feedback at the user's granularity. We study this interface in the context of undergraduate mathematics education by developing LeanSide, a formally verified co-reasoning system, which auto-formalizes student reasoning into Lean and informalizes verifier output into understandable feedback. We conducted user studies through classroom deployment and analyzed which system properties helped students make progress and which caused them to get stuck. We use these findings to derive design implications for using a formally verified backend in human-AI co-reasoning systems.

cs.HC↗