SearcharxivSearch

arXiv subjects

Michaela Benk

Publications and source records attributed to Michaela Benk.

7 recordsLinked to original sources

Ranking Companion: A Visual Analytics Approach to Item-Based Ranking with Hybrid Item Selection

Personalizing item ranking creation is a challenging task, especially when users lack knowledge of data attributes or the ability to express and formalize their attribute preferences. Item-based ranking creation is an approach allowing users to directly externalize preferences through known-item judgments rather than attribute-based scoring. However, a core challenge of item-based ranking is identifying and selecting representative candidate items for externalizing preferences. Existing approaches rely on singular item-selection methods, limiting flexibility and user control. To address this challenge, we present Ranking Companion, a visual analytics approach for item-based ranking that combines model-driven active learning with human-driven item-selection methods. By drawing from six complementary item-selection methods, users can externalize listwise preferences based on selected candidate items, while an iterative machine learning process with a ranking model calculates ranking results, presented to users alongside explanations for interpretation. We evaluated Ranking Companion in a formative user study with 10 participants, in which participants used each item-selection method across three iterations, revealing tradeoffs in perceived ranking quality across accuracy, diversity, novelty, transparency, control, and satisfaction. Ranking Companion contributes a unified interactive item selection space and provides preliminary empirical guidance toward the hybrid use of multiple complementary item-selection methods in personalized item-based ranking creation.

cs.HC

Same Performance, Hidden Bias: Evaluating Hypothesis- and Recommendation-Driven AI

The HCI community commonly evaluates decision support systems based on whether they improve task performance or promote appropriate user reliance. In this work, we look beyond decision outcomes to examine the process through which users develop decision-making strategies. Through a web-based experiment (N = 290) comparing recommendation-driven and hypothesis-driven interaction designs, and using Signal Detection Theory as a theoretical framework, we show that even when performance remains identical, recommendation-driven designs lower participants' thresholds for sufficient evidence and introduce a "hidden bias" in their judgments, resulting in a shifted distribution of errors. Furthermore, we find that experts are just as susceptible to these systemic shifts as novices. We conclude by advocating for a shift in focus: prioritizing decision processes and the preservation of stable evidence standards over performance and reliance alone.

cs.HC

Twenty-Four Years of Empirical Research on Trust in AI: A Bibliometric Review of Trends, Overlooked Issues, and Future Directions

Trust is widely regarded as a critical component to building artificial intelligence (AI) systems that people will use and safely rely upon. As research in this area continues to evolve, it becomes imperative that the research community synchronizes its empirical efforts and aligns on the path toward effective knowledge creation. To lay the groundwork toward achieving this objective, we performed a comprehensive bibliometric analysis, supplemented with a qualitative content analysis of over two decades of empirical research measuring trust in AI, comprising 1'156 core articles and 36'306 cited articles across multiple disciplines. Our analysis reveals several "elephants in the room" pertaining to missing perspectives in global discussions on trust in AI, a lack of contextualized theoretical models and a reliance on exploratory methodologies. We highlight strategies for the empirical research community that are aimed at fostering an in-depth understanding of trust in AI.

cs.HC

Certification Labels for Trustworthy AI: Insights From an Empirical Mixed-Method Study

Auditing plays a pivotal role in the development of trustworthy AI. However, current research primarily focuses on creating auditable AI documentation, which is intended for regulators and experts rather than end-users affected by AI decisions. How to communicate to members of the public that an AI has been audited and considered trustworthy remains an open challenge. This study empirically investigated certification labels as a promising solution. Through interviews (N = 12) and a census-representative survey (N = 302), we investigated end-users' attitudes toward certification labels and their effectiveness in communicating trustworthiness in low- and high-stakes AI scenarios. Based on the survey results, we demonstrate that labels can significantly increase end-users' trust and willingness to use AI in both low- and high-stakes scenarios. However, end-users' preferences for certification labels and their effect on trust and willingness to use AI were more pronounced in high-stake scenarios. Qualitative content analysis of the interviews revealed opportunities and limitations of certification labels, as well as facilitators and inhibitors for the effective use of labels in the context of AI. For example, while certification labels can mitigate data-related concerns expressed by end-users (e.g., privacy and data protection), other concerns (e.g., model performance) are more challenging to address. Our study provides valuable insights and recommendations for designing and implementing certification labels as a promising constituent within the trustworthy AI ecosystem.

cs.CY

"Is It My Turn?" Assessing Teamwork and Taskwork in Collaborative Immersive Analytics

Immersive analytics has the potential to promote collaboration in machine learning (ML). This is desired due to the specific characteristics of ML modeling in practice, namely the complexity of ML, the interdisciplinary approach in industry, and the need for ML interpretability. In this work, we introduce an augmented reality-based system for collaborative immersive analytics that is designed to support ML modeling in interdisciplinary teams. We conduct a user study to examine how collaboration unfolds when users with different professional backgrounds and levels of ML knowledge interact in solving different ML tasks. Specifically, we use the pair analytics methodology and performance assessments to assess collaboration and explore their interactions with each other and the system. Based on this, we provide qualitative and quantitative results on both teamwork and taskwork during collaboration. Our results show how our system elicits sustained collaboration as measured along six distinct dimensions. We finally make recommendations how immersive systems should be designed to elicit sustained collaboration in ML modeling.

cs.HC

Creative Uses of AI Systems and their Explanations: A Case Study from Insurance

Recent works have recognized the need for human-centered perspectives when designing and evaluating human-AI interactions and explainable AI methods. Yet, current approaches fall short at intercepting and managing unexpected user behavior resulting from the interaction with AI systems and explainability methods of different stake-holder groups. In this work, we explore the use of AI and explainability methods in the insurance domain. In an qualitative case study with participants with different roles and professional backgrounds, we show that AI and explainability methods are used in creative ways in daily workflows, resulting in a divergence between their intended and actual use. Finally, we discuss some recommendations for the design of human-AI interactions and explainable AI methods to manage the risks and harness the potential of unexpected user behavior.

cs.HC

The Value of Measuring Trust in AI - A Socio-Technical System Perspective

Building trust in AI-based systems is deemed critical for their adoption and appropriate use. Recent research has thus attempted to evaluate how various attributes of these systems affect user trust. However, limitations regarding the definition and measurement of trust in AI have hampered progress in the field, leading to results that are inconsistent or difficult to compare. In this work, we provide an overview of the main limitations in defining and measuring trust in AI. We focus on the attempt of giving trust in AI a numerical value and its utility in informing the design of real-world human-AI interactions. Taking a socio-technical system perspective on AI, we explore two distinct approaches to tackle these challenges. We provide actionable recommendations on how these approaches can be implemented in practice and inform the design of human-AI interactions. We thereby aim to provide a starting point for researchers and designers to re-evaluate the current focus on trust in AI, improving the alignment between what empirical research paradigms may offer and the expectations of real-world human-AI interactions.

cs.HC