paper-with-me

홈 › Papers

Listening without Looking: Modality Bias in Audio-Visual Captioning

2025-10-28 · Yuchi Ishikawa, Toranosuke Manabe, Tatsuya Komatsu, Yoshimitsu Aoki arxiv

Audio-visual captioning aims to generate holistic scene descriptions by jointly modeling sound and vision. While recent methods have improved performance through sophisticated modality fusion, it remains unclear to what extent the two modalities are complementary in current audio-visual captioning models and how robust these models are when one modality is degraded. We address these questions by conducting systematic modality robustness tests on LAVCap, a state-of-the-art audio-visual captioning model, in which we selectively suppress or corrupt the audio or visual streams to quantify sensitivity and complementarity. The analysis reveals a pronounced bias toward the audio stream in LAVCap. To evaluate how balanced audio-visual captioning models are in their use of both modalities, we augment AudioCaps with textual annotations that jointly describe the audio and visual streams, yielding the AudioVisualCaps dataset. In our experiments, we report LAVCap baseline results on AudioVisualCaps. We also evaluate the model under modality robustness tests on AudioVisualCaps and the results indicate that LAVCap trained on AudioVisualCaps exhibits less modality bias than when trained on AudioCaps.

📄 PDF Abstract BibTeX arXiv:2510.24024

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Read, Look or Listen? What's Needed for Solving a Multimodal Dataset

2023-07-06 · Netta Madvil, Yonatan Bitton, Roy Schwartz

The prevalence of large-scale multimodal datasets presents unique challenges in assessing dataset quality. We propose a two-step method to analyze multimodal datasets, which leverages a small seed of human annotation to …

Question AnsweringSpeaker IdentificationVideo Question Answering

Looking and Listening Inside and Outside: Multimodal Artificial Intelligence Systems for Driver Safety Assessment and Intelligent Vehicle Decision-Making

2026-02-07 · Ross Greer, Laura Fleig, Maitrayee Keskar, Erika Maquiling 외 arxiv

The looking-in-looking-out (LILO) framework has enabled intelligent vehicle applications that understand both the outside scene and the driver state to improve safety outcomes, with examples in smart airbag deployment, t…

Scene Understanding

AV Taris: Online Audio-Visual Speech Recognition

2020-12-14 · George Sterpu, Naomi Harte

In recent years, Automatic Speech Recognition (ASR) technology has approached human-level performance on conversational speech under relatively clean listening conditions. In more demanding situations involving distant m…

Action DetectionActivity DetectionAudio-Visual Speech RecognitionAutomatic Speech Recognition+4

Audio-Visual Neural Syntax Acquisition

2023-10-11 · Cheng-I Jeff Lai, Freda Shi, Puyuan Peng, Yoon Kim 외

We study phrase structure induction from visually-grounded speech. The core idea is to first segment the speech waveform into sequences of word segments, and subsequently induce phrase structure using the inferred segmen…

Language Acquisition

Look, Listen and Learn

2017-05-23 · ICCV 2017 10 · Relja Arandjelović, Andrew Zisserman

We consider the question: what can be learnt by looking at and listening to a large number of unlabelled videos? There is a valuable, but so far untapped, source of information contained in the video itself -- the corres…

Audio ClassificationGeneral ClassificationSound Classification