SearcharxivSearch

arXiv subjects

Gerd Kortuem

Publications and source records attributed to Gerd Kortuem.

8 recordsLinked to original sources

MarkupLens: Balancing Computer Vision Assistance and Control in Professional Video Annotation for Video-Based Design Tasks

Video-Based Design (VBD) uses video as a primary medium for analyzing user interactions, prototyping, and generating design insights. However, current VBD workflows are constrained by labor-intensive, inconsistent manual annotations that fragment attention and delay insights. Computer Vision (CV)-powered automatic annotation offers opportunities to reduce manual effort while supporting higher-level interpretation. This paper investigates human-AI collaboration in video analysis by examining how different levels of automated support shape user experience in VBD. We developed MarkupLens, a CV-assisted annotation platform, and conducted a between-subjects eye-tracking study with 36 designers in an urban VBD case. We compared three levels of automation: no support, partial support, and full support, and found that higher levels improved annotation quality, reduced cognitive load, and interestingly, enriched reflection. Our insights on automation levels inform adjustable autonomy and mixed-initiative system design beyond VBD tasks.

cs.HC

Understanding Mental Models of Generative Conversational Search and The Effect of Interface Transparency

The experience and adoption of conversational search is tied to the accuracy and completeness of users' mental models -- their internal frameworks for understanding and predicting system behaviour. Thus, understanding these models can reveal areas for design interventions. Transparency is one such intervention which can improve system interpretability and enable mental model alignment. While past research has explored mental models of search engines, those of generative conversational search remain underexplored, even while the popularity of these systems soars. To address this, we conducted a study with 16 participants, who performed 4 search tasks using 4 conversational interfaces of varying transparency levels. Our analysis revealed that most user mental models were too abstract to support users in explaining individual search instances. These results suggest that 1) mental models may pose a barrier to appropriate trust in conversational search, and 2) hybrid web-conversational search is a promising novel direction for future search interface design.

cs.HC

"A Great Start, But...": Evaluating LLM-Generated Mind Maps for Information Mapping in Video-Based Design

Extracting concepts and understanding relationships from videos is essential in Video-Based Design (VBD), where videos serve as a primary medium for exploration but require significant effort in managing meta-information. Mind maps, with their ability to visually organize complex data, offer a promising approach for structuring and analysing video content. Recent advancements in Large Language Models (LLMs) provide new opportunities for meta-information processing and visual understanding in VBD, yet their application remains underexplored. This study recruited 28 VBD practitioners to investigate the use of prompt-tuned LLMs for generating mind maps from ethnographic videos. Comparing LLM-generated mind maps with those created by professional designers, we evaluated rated scores, design effectiveness, and user experience across two contexts. Findings reveal that LLMs effectively capture central concepts but struggle with hierarchical organization and contextual grounding. We discuss trust, customization, and workflow integration as key factors to guide future research on LLM-supported information mapping in VBD.

cs.HC

Tangi: a Tool to Create Tangible Artifacts for Sharing Insights from 360$^\circ$ Video

Designers often engage with video to gain rich, temporal insights about the context of users, collaboratively analyzing it to gather ideas, challenge assumptions, and foster empathy. To capture the full visual context of users and their situations, designers are adopting 360$^\circ$ video, providing richer, more multi-layered insights. Unfortunately, the spherical nature of 360$^\circ$ video means designers cannot create tangible video artifacts such as storyboards for collaborative analysis. To overcome this limitation, we created Tangi, a web-based tool that converts 360$^\circ$ images into tangible 360$^\circ$ video artifacts, that enable designers to embody and share their insights. Our evaluation with nine experienced designers demonstrates that the artifacts Tangi creates enable tangible interactions found in collaborative workshops and introduce two new capabilities: spatial orientation within 360$^\circ$ environments and linking specific details to the broader 360$^\circ$ context. Since Tangi is an open-source tool, designers can immediately leverage 360$^\circ$ video in collaborative workshops.

cs.HC

DesignMinds: Enhancing Video-Based Design Ideation with Vision-Language Model and Context-Injected Large Language Model

Ideation is a critical component of video-based design (VBD), where videos serve as the primary medium for design exploration and inspiration. The emergence of generative AI offers considerable potential to enhance this process by streamlining video analysis and facilitating idea generation. In this paper, we present DesignMinds, a prototype that integrates a state-of-the-art Vision-Language Model (VLM) with a context-enhanced Large Language Model (LLM) to support ideation in VBD. To evaluate DesignMinds, we conducted a between-subject study with 35 design practitioners, comparing its performance to a baseline condition. Our results demonstrate that DesignMinds significantly enhances the flexibility and originality of ideation, while also increasing task engagement. Importantly, the introduction of this technology did not negatively impact user experience, technology acceptance, or usability.

cs.HC

Sphere Window: Challenges and Opportunities of 360° Video in Collaborative Design Workshops

The increased ubiquity of 360° video presents a unique opportunity for designers to deeply engage with the world of users by capturing the complete visual context. However, the opportunities and challenges 360° video introduces for video design ethnography is unclear. This study investigates this gap through 16 workshops in which experienced designers engaged with 360° video. Our analysis shows that while 360° video enhances designers' ability to explore and understand user contexts, it also complicates the process of sharing insights. To address this challenge, we present two opportunities to support the use of 360° video by designers - the creation of designerly 360° video annotation tools, and 360° ``screenshots'' - in order to enable designers to leverage the complete context of 360° video for user research.

cs.HC

Contestable Camera Cars: A Speculative Design Exploration of Public AI That Is Open and Responsive to Dispute

Local governments increasingly use artificial intelligence (AI) for automated decision-making. Contestability, making systems responsive to dispute, is a way to ensure they respect human rights to autonomy and dignity. We investigate the design of public urban AI systems for contestability through the example of camera cars: human-driven vehicles equipped with image sensors. Applying a provisional framework for contestable AI, we use speculative design to create a concept video of a contestable camera car. Using this concept video, we then conduct semi-structured interviews with 17 civil servants who work with AI employed by a large northwestern European city. The resulting data is analyzed using reflexive thematic analysis to identify the main challenges facing the implementation of contestability in public AI. We describe how civic participation faces issues of representation, public AI systems should integrate with existing democratic practices, and cities must expand capacities for responsible AI development and operation.

cs.HC

Micro-Navigation for Urban Bus Passengers: Using the Internet of Things to Improve the Public Transport Experience

Public bus services are widely deployed in cities around the world because they provide cost-effective and economic public transportation. However, from a passenger point of view urban bus systems can be complex and difficult to navigate, especially for disadvantaged users, i.e. tourists, novice users, older people, and people with impaired cognitive or physical abilities. We present Urban Bus Navigator (UBN), a reality-aware urban navigation system for bus passengers with the ability to recognize and track the physical public transport infrastructure such as buses. Unlike traditional location-aware mobile transport applications, UBN acts as a true navigation assistant for public transport users. Insights from a six-month long trial in Madrid indicate that UBN removes barriers for public transport usage and has a positive impact on how people feel about public transport journeys.

cs.HC