SearcharxivSearch

arXiv · 2602.01390

Toward Scalable Audio Description Quality Control: A Workflow for Evaluating Human and VLM Raters

Abstract

Digital video is central to communication, education, and entertainment, but without audio description (AD), blind and low-vision users are excluded. While crowdsourced platforms and vision-language models (VLMs) expand AD production, quality is rarely checked systematically. Existing evaluations rely on NLP metrics and short-clip guidelines, leaving open the question of how to assess long-form AD quality at scale. To address this, we developed a methodological workflow using Item Response Theory to evaluate VLM and human rater proficiency against expert-established ground truth. Evaluations were based on a six-dimensional framework, grounded in professional guidelines and shaped by insights from our accessibility experts and blind consultants. Findings suggest that top-performing VLMs can approximate ground-truth ratings at levels comparable to human raters. However, qualitative analysis reveals that VLM reasoning is less reliable and actionable than that of human respondents. These insights underscore the potential of hybrid evaluation systems that leverage VLMs alongside human oversight, offering a path toward scalable AD quality control.

Explore related subjects

Keep this discovery

BibTeXRIS

Lana Do, Gio Jung, Juvenal Francisco Barajas, Andrew Taylor Scott, Shasta Ihorn, Alexander Mario Blum, Vassilis Athitsos, Ilmi Yoon. 2026-08-30. Toward Scalable Audio Description Quality Control: A Workflow for Evaluating Human and VLM Raters. https://arxiv.org/abs/2602.01390

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related discoveries

Experts Disagree on How to Fight AI Disinformation, but Agree That Health and Politics Need Different Solutions

When 54 international experts assessed AI-generated disinformation threats, they revealed a surprising pattern: while video deepfakes received the highest average threat ratings in the political domain (M = 6.31/7), the pattern differed in the health domain, where AI-generated text received the highest average rating (M = 5.80). Experts also diverge on what to do: government regulation drew both the most "most effective" (30%) and the most "least effective" (15%) votes, though rating distributions were contested rather than polarized, indicating disagreement over priorities rather than over efficacy. These findings offer an initial expert map of an AI-disinformation landscape that is still rapidly forming.

cs.CY

Generative AI Alignment with Hinduism's Theological Plurality and Sacred Representation

Generative AI systems are increasingly used to answer personal questions and mediate everyday practices, including religion. However, existing discussions around AI alignment and ethics have largely centered secular, Western, and Abrahamic assumptions about religion, offering limited attention to other faith-based traditions. In this paper, we examine how Hindu users engage with generative AI systems in relation to their religious knowledge, belief, and practice. Drawing on 15 semi-structured interviews with Bangladeshi Hindu participants, we analyze how users interpret AI-generated religious representations, scriptural explanations, devotional interactions, and synthetic religious media. We found that AI can be both accessible and ethically troubling. While AI supported scriptural inquiry, devotional visualization, and religious storytelling, our study also identified concerns about theological flattening, cultural misrepresentation, devotional manipulation, and the simulation of sacred presence and authority. We conclude by arguing that religious alignment in generative AI requires interpretive alignment: systems that disclose their limits, preserve plurality, and avoid simulating sacred authority and sycophantic personalization.

cs.AI

Measuring the "Interaction Gap" in Drama Therapy with AI

Generative AI is increasingly being introduced into expressive arts therapy, where it is often credited with offering a non-judgmental environment that supports psychological safety. Existing HCI work has largely positioned AI as a co-creative material or as a bridge/mediator into human-led care. This paper explores a different position. When a patient performs the same drama therapy task with an AI partner and with a human partner, the resulting self-presentations tend to differ in patterned ways. We propose treating this difference, the Interaction Gap, as a diagnostic lens within drama therapy. Rather than asking which context elicits a truer self, the lens reads the difference between the two performances as information about the social pressures shaping self-expression in each context. We sketch a starting point for task design and measurement signals grounded in drama therapy's existing use of role and aesthetic distance, and raise provocations for workshop discussion: the observer effect and privacy paradox that measurement introduces, and the question of whose lens the gap is.

cs.HC