Searcharxiv⌕ Search

arXiv · 2609.35197

Alignment Games: A Framework for Conceptual Repair in Human-AI Collaboration

Abstract

The meaning of a concept in use is shaped by the situation, task, goals, and prior knowledge. For example, a request to make a poster "visually appealing for a five-year-old" might evoke bright colors and cartoon imagery for one collaborator, but less text, bold shapes, and visual simplicity for another. We call such task-relevant differences conceptual misalignment. We introduce Alignment Games, a framework for making these differences visible and repairable during human-AI interaction. Drawing on theories of situated conceptualization, we characterize task-specific conceptual frames in terms of relevant attributes, values, relations, constraints, and priorities. We then define alignment moves that intervene on the situation, the reasoning used to interpret it, or the resulting frame. Through examples from educational content generation, creative coding, and argumentative writing, we show how these moves can be composed into repair sequences and derive design principles for supporting task-sufficient conceptual alignment at runtime.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hari Subramonyam, Maneesh Agrawala, Sean Follmer. 2026-09-28. Alignment Games: A Framework for Conceptual Repair in Human-AI Collaboration. https://arxiv.org/abs/2609.35197

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

HAGI++: Head-Assisted Gaze Imputation and Generation

Mobile eye-tracking is crucial for capturing human visual attention in real-world and XR settings, supporting research and human-computer interaction. Yet blinks, pupil-detection errors and lighting changes create missing values that hinder gaze analysis. We present HAGI++, a multi-modal diffusion-based imputation method that, for the first time, leverages integrated head-orientation sensors to exploit the natural correlation between head and eye movements. Using a transformer-based diffusion model, it learns cross-modal dependencies between eye and head data and can additionally incorporate wrist/hand motion when such wearable signals are available. Evaluations on the large-scale Nymeria, Ego-Exo4D and HOT3D datasets show that HAGI++ consistently outperforms traditional interpolation and deep-learning time-series imputation baselines. Statistical analysis confirms that its gaze-velocity distributions closely match real human behaviour, yielding realistic imputations. Even when 100% of gaze data are missing (pure gaze generation), HAGI++ exceeds methods that rely on the visual inputs and the methods rely on full-body motion capture by incorporating wrist motion from commercial wearables. Our approach enables more complete, accurate eye-gaze recordings in real-world contexts, enhancing gaze-based analysis and interaction across many applications. Our code is available at https://git.cai.simtech.uni-stuttgart.de/public-projects/HAGI

cs.HC↗

Deco: Extending Cherished Physical Objects into AI Companion Agents through Dual Embodiment

Physical objects (e.g., plush toys) can transcend materiality to become emotional anchors and provide companionship. However, these bonds remain one-sided because most physical objects cannot reciprocate. AI companions offer responsiveness and personalization, but typically entail building bonds from scratch. We investigated how AI companions might inherit and extend users' existing bonds with physical objects. A formative study (N=9) informed four design principles (Faithful Identity, Calibrated Agency, Ambient Presence, Reciprocal Memory), shaping our Dual-Embodiment Companion Framework. We instantiated it as Deco to create digital embodiments of physical companions. In a within-subjects lab study (N=25), Deco was rated higher than a personalized digital-only companion on six companion-related measures (all p<.01). A subsequent seven-day field deployment (N=17) showed sustained engagement, higher post-deployment well-being (p=.040), and three key relational patterns: digital activities retroactively vitalized physical objects, bond deepening centered on emotional engagement depth, and participants sustained bonds while navigating companions' AI nature. Dual embodiment offers a promising framework for revitalizing physical objects with AI agents.

cs.HC↗

OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments

Current benchmarks for graphical user interface (GUI) agents predominantly rely on static screenshots. However, real-world smartphone interaction routinely requires agents to process transient audio cues and temporal video dynamics that are tightly coupled with the moment of action. To bridge this gap, we introduce OmniGUI, the first step-level benchmark designed to evaluate GUI agents in omni-modal smartphone environments. OmniGUI provides continuous, interleaved multimodal inputs comprising static images, synchronous audio, and video clips at every action step. The dataset encompasses 709 expert-demonstrated episodes (2,579 action steps) across 29 applications, systematically annotated with objective multimodal dependency levels. Because dedicated omni-modal GUI agent frameworks are currently in their nascent stage, we select foundational omni-modal models capable of natively processing interleaved inputs to serve as agent proxies for our initial baselines. Our empirical evaluation reveals that while current models exhibit competency on visually static tasks, their action prediction performance degrades significantly in environments requiring synchronous temporal and auditory signals. Furthermore, ablation studies isolate specific operational bottlenecks, notably cross-modal interference when processing task-irrelevant environmental noise. The complete dataset, evaluation pipeline, and baseline prompts are provided in the supplementary material. Project page: https://omni-gui.github.io.

cs.HC↗