arXiv · 2602.21337
A Benchmark to Assess Common Ground in Human-AI Collaboration
Abstract
AI is becoming increasingly integrated into everyday life, both in professional work environments and in leisure and entertainment contexts. This integration requires AI to move beyond acting as an assistant for informational or transactional tasks toward a genuine collaborative partner. Effective collaboration, whether between humans or between humans and AI, depends on establishing and maintaining common ground: shared beliefs, assumptions, goals, and situational awareness that enable coordinated action and efficient repair of misunderstandings. While common ground is a central concept in human collaboration, it has received limited attention in studies of human-AI collaboration. In this paper, we introduce a new benchmark grounded in theories and empirical studies of human-human collaboration. The benchmark is based on a collaborative puzzle task that requires iterative interaction, joint action, referential coordination, and repair under varying conditions of situation awareness. We validate the benchmark through a confirmatory user study in which human participants collaborate with an AI to solve the task. The results show that the benchmark reproduces established theoretical and empirical findings from human-human collaboration, while also revealing clear divergences in human-AI interaction.
Explore related subjects
Keep this discovery
Christian Poelitz, Finale Doshi-Velez, Siân Lindley. 2026-02-24. A Benchmark to Assess Common Ground in Human-AI Collaboration. https://arxiv.org/abs/2602.21337
Cite the original work for its findings. Save a collection to share your selection of sources.