paper-with-me

홈 › Papers

Toward Scalable Audio Description Quality Control: A Workflow for Evaluating Human and VLM Raters

2026-02-01 · Lana Do, Gio Jung, Juvenal Francisco Barajas, Andrew Taylor Scott, Shasta Ihorn, Alexander Mario Blum, Vassilis Athitsos, Ilmi Yoon arxiv

Digital video is central to communication, education, and entertainment, but without audio description (AD), blind and low-vision users are excluded. While crowdsourced platforms and vision-language models (VLMs) expand AD production, quality is rarely checked systematically. Existing evaluations rely on NLP metrics and short-clip guidelines, leaving open the question of how to assess long-form AD quality at scale. To address this, we developed a methodological workflow using Item Response Theory to evaluate VLM and human rater proficiency against expert-established ground truth. Evaluations were based on a six-dimensional framework, grounded in professional guidelines and shaped by insights from our accessibility experts and blind consultants. Findings suggest that top-performing VLMs can approximate ground-truth ratings at levels comparable to human raters. However, qualitative analysis reveals that VLM reasoning is less reliable and actionable than that of human respondents. These insights underscore the potential of hybrid evaluation systems that leverage VLMs alongside human oversight, offering a path toward scalable AD quality control.

📄 PDF Abstract BibTeX arXiv:2602.01390

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Aligning Audio-Visual Joint Representations with an Agentic Workflow

2024-10-30 · Shentong Mo, Yibing Song

Visual content and accompanied audio signals naturally formulate a joint representation to improve audio-visual (AV) related applications. While studies develop various AV representation learning frameworks, the importan…

Representation Learning

AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions

2024-09-19 · Yuanyuan Wang, Hangting Chen, Dongchao Yang, Zhiyong Wu 외

Current Text-to-audio (TTA) models mainly use coarse text descriptions as inputs to generate audio, which hinders models from generating audio with fine-grained control of content and style. Some studies try to improve t…

Audio Generation

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment

2024-12-19 · Riccardo Fosco Gramaccioni, Christian Marinoni, Emilian Postolache, Marco Comunità 외

Traditional sound design workflows rely on manual alignment of audio events to visual cues, as in Foley sound design, where everyday actions like footsteps or object interactions are recreated to match the on-screen moti…

Audio Generation

CANVAS: Captioning Art with Narrative Visual-Audio AI Systems

2026-04-30 · Vignesh Nagarajan arxiv

Visual art remains largely inaccessible to blind and low-vision (BLV) audiences due to brief or absent alt-text, which rarely conveys the sensory, spatial, or emotional qualities of an artwork. This study presents an aut…

DegDiT: Controllable Audio Generation with Dynamic Event Graph Guided Diffusion Transformer

2025-08-19 · Yisu Liu, Chenxing Li, Wanqian Zhang, Wenfu Wang 외 arxiv

Controllable text-to-audio generation aims to synthesize audio from textual descriptions while satisfying user-specified constraints, including event types, temporal sequences, and onset and offset timestamps. This enabl…

Temporal SequencesAudio Generation