paper-with-me

Papers

FocusedAD: Character-centric Movie Audio Description

2025-04-16 · Xiaojun Ye, Chun Wang, Yiren Song, Sheng Zhou, Liangcheng Li, Jiajun Bu

Movie Audio Description (AD) aims to narrate visual content during dialogue-free segments, particularly benefiting blind and visually impaired (BVI) audiences. Compared with general video captioning, AD demands plot-relevant narration with explicit character name references, posing unique challenges in movie understanding.To identify active main characters and focus on storyline-relevant regions, we propose FocusedAD, a novel framework that delivers character-centric movie audio descriptions. It includes: (i) a Character Perception Module(CPM) for tracking character regions and linking them to names; (ii) a Dynamic Prior Module(DPM) that injects contextual cues from prior ADs and subtitles via learnable soft prompts; and (iii) a Focused Caption Module(FCM) that generates narrations enriched with plot-relevant details and named characters. To overcome limitations in character identification, we also introduce an automated pipeline for building character query banks. FocusedAD achieves state-of-the-art performance on multiple benchmarks, including strong zero-shot results on MAD-eval-Named and our newly proposed Cinepile-AD dataset. Code and data will be released at https://github.com/Thorin215/FocusedAD .

📄 PDF Abstract BibTeX arXiv:2504.12157

Code (1)

thorin215/focusedad 공식 구현 pytorch

Tasks

Video Captioning

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Character-Centric Understanding of Animated Movies

2025-09-15 · Zhongrui Gui, Junyu Xie, Tengda Han, Weidi Xie 외 arxiv

Animated movies are captivating for their unique character designs and imaginative storytelling, yet they pose significant challenges for existing recognition systems. Unlike the consistent visual patterns detected by co…

Face Recognition

StoryTeller: Improving Long Video Description through Global Audio-Visual Character Identification

2024-11-11 · Yichen He, Yuan Lin, Jianchao Wu, Hanchong Zhang 외

Existing large vision-language models (LVLMs) are largely limited to processing short, seconds-long videos and struggle with generating coherent descriptions for extended video spanning minutes or more. Long video descri…

Large Language ModelMultimodal Large Language ModelMultiple-choiceVideo Description

Movie Description

2016-05-12 · Anna Rohrbach, Atousa Torabi, Marcus Rohrbach, Niket Tandon 외

Audio Description (AD) provides linguistic descriptions of movies and allows visually impaired people to follow a movie along with their peers. Such descriptions are by design mainly visual and thus naturally form an int…

Benchmarking

AutoAD II: The Sequel -- Who, When, and What in Movie Audio Description

2023-10-10 · Tengda Han, Max Bain, Arsha Nagrani, Gül Varol 외

Audio Description (AD) is the task of generating descriptions of visual content, at suitable time intervals, for the benefit of visually impaired audiences. For movies, this presents notable challenges -- AD must occur o…

Language ModellingText Generation

AutoAD II: The Sequel - Who, When, and What in Movie Audio Description

2023-01-01 · ICCV 2023 1 · Tengda Han, Max Bain, Arsha Nagrani, Gul Varol 외

Audio Description (AD) is the task of generating descriptions of visual content, at suitable time intervals, for the benefit of visually impaired audiences. For movies, this presents notable challenges -- AD must occ…

Language ModellingText Generation