paper-with-me

홈 › Papers

What You See is What You Ask: Evaluating Audio Descriptions

2025-10-01 · Divy Kala, Eshika Khandelwal, Makarand Tapaswi arxiv

Audio descriptions (ADs) narrate important visual details in movies, enabling Blind and Low Vision (BLV) users to understand narratives and appreciate visual details. Existing works in automatic AD generation mostly focus on few-second trimmed clips, and evaluate them by comparing against a single ground-truth reference AD. However, writing ADs is inherently subjective. Through alignment and analysis of two independent AD tracks for the same movies, we quantify the subjectivity in when and whether to describe, and what and how to highlight. Thus, we show that working with trimmed clips is inadequate. We propose ADQA, a QA benchmark that evaluates ADs at the level of few-minute long, coherent video segments, testing whether they would help BLV users understand the story and appreciate visual details. ADQA features visual appreciation (VA) questions about visual facts and narrative understanding (NU) questions based on the plot. Through ADQA, we show that current AD generation methods lag far behind human-authored ADs. We conclude with several recommendations for future work and introduce a public leaderboard for benchmarking.

📄 PDF Abstract BibTeX arXiv:2510.00808

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models

2025-05-12 · Riccardo Passoni, Francesca Ronchini, Luca Comanducci, Romain Serizel 외

Text-to-audio models have recently emerged as a powerful technology for generating sound from textual descriptions. However, their high computational demands raise concerns about energy consumption and environmental impa…

Spoken Moments: Learning Joint Audio-Visual Representations from Video Descriptions

2021-05-10 · CVPR 2021 1 · Mathew Monfort, SouYoung Jin, Alexander Liu, David Harwath 외

When people observe events, they are able to abstract key information and build concise summaries of what is happening. These summaries include contextual and semantic information describing the important high-level deta…

Contrastive LearningRetrievalVideo Understanding

What You Hear Is What You See: Audio Quality Metrics From Image Quality Metrics

2023-05-19 · Tashi Namgyal, Alexander Hepburn, Raul Santos-Rodriguez, Valero Laparra 외

In this study, we investigate the feasibility of utilizing state-of-the-art image perceptual metrics for evaluating audio signals by representing them as spectrograms. The encouraging outcome of the proposed approach is …

Learning Spatially-Aware Language and Audio Embeddings

2024-09-17 · Bhavika Devnani, Skyler Seto, Zakaria Aldeneh, Alessandro Toso 외

Humans can picture a sound scene given an imprecise natural language description. For example, it is easy to imagine an acoustic environment given a phrase like "the lion roar came from right behind me!". For a machine t…

AttributeContrastive LearningPositionSemantic Retrieval+1

Movie Description

2016-05-12 · Anna Rohrbach, Atousa Torabi, Marcus Rohrbach, Niket Tandon 외

Audio Description (AD) provides linguistic descriptions of movies and allows visually impaired people to follow a movie along with their peers. Such descriptions are by design mainly visual and thus naturally form an int…

Benchmarking