paper-with-me

홈 › Papers

DistinctAD: Distinctive Audio Description Generation in Contexts

2024-11-27 · CVPR 2025 1 · Bo Fang, Wenhao Wu, Qiangqiang Wu, Yuxin Song, Antoni B. Chan

Audio Descriptions (ADs) aim to provide a narration of a movie in text form, describing non-dialogue-related narratives, such as characters, actions, or scene establishment. Automatic generation of ADs remains challenging due to: i) the domain gap between movie-AD data and existing data used to train vision-language models, and ii) the issue of contextual redundancy arising from highly similar neighboring visual clips in a long movie. In this work, we propose DistinctAD, a novel two-stage framework for generating ADs that emphasize distinctiveness to produce better narratives. To address the domain gap, we introduce a CLIP-AD adaptation strategy that does not require additional AD corpora, enabling more effective alignment between movie and AD modalities at both global and fine-grained levels. In Stage-II, DistinctAD incorporates two key innovations: (i) a Contextual Expectation-Maximization Attention (EMA) module that reduces redundancy by extracting common bases from consecutive video clips, and (ii) an explicit distinctive word prediction loss that filters out repeated words in the context, ensuring the prediction of unique terms specific to the current AD. Comprehensive evaluations on MAD-Eval, CMD-AD, and TV-AD benchmarks demonstrate the superiority of DistinctAD, with the model consistently outperforming baselines, particularly in Recall@k/N, highlighting its effectiveness in producing high-quality, distinctive ADs.

📄 PDF Abstract BibTeX arXiv:2411.18180

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Generating event descriptions under syntactic and semantic constraints

2024-12-24 · Angela Cao, Faye Holt, Jonas Chan, Stephanie Richter 외

With the goal of supporting scalable lexical semantic annotation, analysis, and theorizing, we conduct a comprehensive evaluation of different methods for generating event descriptions under both syntactic constraints --…

Language ModelingLanguage Modelling

MetaBGM: Dynamic Soundtrack Transformation For Continuous Multi-Scene Experiences With Ambient Awareness And Personalization

2024-09-05 · Haoxuan Liu, ZiHao Wang, HaoRong Hong, Youwei Feng 외

This paper introduces MetaBGM, a groundbreaking framework for generating background music that adapts to dynamic scenes and real-time user interactions. We define multi-scene as variations in environmental contexts, such…

Audio Generation

Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation

2024-10-14 · Peiwen Sun, Sitong Cheng, Xiangtai Li, Zhen Ye 외

Recently, diffusion models have achieved great success in mono-channel audio generation. However, when it comes to stereo audio generation, the soundscapes often have a complex scene of multiple objects and directions. C…

Audio Generationmultimodal generation

Visually-Aware Audio Captioning With Adaptive Audio-Visual Attention

2022-10-28 · Xubo Liu, Qiushi Huang, Xinhao Mei, Haohe Liu 외

Audio captioning aims to generate text descriptions of audio clips. In the real world, many objects produce similar sounds. How to accurately recognize ambiguous sounds is a major challenge for audio captioning. In this …

AudioCapsAudio captioningMachine Translation

On The Open Prompt Challenge In Conditional Audio Generation

2023-11-01 · Ernie Chang, Sidd Srinivasan, Mahi Luthra, Pin-Jie Lin 외

Text-to-audio generation (TTA) produces audio from a text description, learning from pairs of audio samples and hand-annotated text. However, commercializing audio generation is challenging as user-input prompts are ofte…

Audio Generation