paper-with-me

Papers

CAD -- Contextual Multi-modal Alignment for Dynamic AVQA

2023-10-25 · Asmar Nadeem, Adrian Hilton, Robert Dawes, Graham Thomas, Armin Mustafa

In the context of Audio Visual Question Answering (AVQA) tasks, the audio visual modalities could be learnt on three levels: 1) Spatial, 2) Temporal, and 3) Semantic. Existing AVQA methods suffer from two major shortcomings; the audio-visual (AV) information passing through the network isn't aligned on Spatial and Temporal levels; and, inter-modal (audio and visual) Semantic information is often not balanced within a context; this results in poor performance. In this paper, we propose a novel end-to-end Contextual Multi-modal Alignment (CAD) network that addresses the challenges in AVQA methods by i) introducing a parameter-free stochastic Contextual block that ensures robust audio and visual alignment on the Spatial level; ii) proposing a pre-training technique for dynamic audio and visual alignment on Temporal level in a self-supervised setting, and iii) introducing a cross-attention mechanism to balance audio and visual information on Semantic level. The proposed novel CAD network improves the overall performance over the state-of-the-art methods on average by 9.4% on the MUSIC-AVQA dataset. We also demonstrate that our proposed contributions to AVQA can be added to the existing methods to improve their performance without additional complexity requirements.

📄 PDF Abstract BibTeX arXiv:2310.16754

Code (0)

등록된 구현이 없습니다.

Tasks

Audio-visual Question AnsweringAudio-Visual Question Answering (AVQA)Question AnsweringVisual Question Answering

Similar Papers 제목 키워드 기반

MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding

2025-05-26 · Jeonghun Baek, Kazuki Egashira, Shota Onohara, Atsuyuki Miyai 외

Manga, or Japanese comics, is a richly multimodal narrative form that blends images and text in complex ways. Teaching large multimodal models (LMMs) to understand such narratives at a human-like level could help manga c…

Question AnsweringVisual Question Answering

Learning to Answer Questions in Dynamic Audio-Visual Scenarios

2022-03-26 · CVPR 2022 1 · Guangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu 외

In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in videos. The problem requires comprehensive …

audio-visual learningAudio-visual Question AnsweringAudio-Visual Question Answering (AVQA)AUDIO-VISUAL QUESTION ANSWERING (MUSIC-AVQA-v2.0)+4

AVQACL: A Novel Benchmark for Audio-Visual Question Answering Continual Learning

2025-01-01 · CVPR 2025 1 · Kaixuan Wu, Xinde Li, Xinling Li, Chuanfei Hu 외

In this paper, a novel benchmark for audio-visual question answering continual learning (AVQACL) is introduced, aiming to study fine-grained scene understanding and spatial-temporal reasoning in videos under a contin…

Audio-visual Question AnsweringContinual LearningKnowledge DistillationQuestion Answering+2

Music's Multimodal Complexity in AVQA: Why We Need More than General Multimodal LLMs

2025-05-27 · Wenhao You, Xingjian Diao, Chunhui Zhang, Keyi Kong 외

While recent Multimodal Large Language Models exhibit impressive capabilities for general multimodal tasks, specialized domains like music necessitate tailored approaches. Music Audio-Visual Question Answering (Music AVQ…

Audio-visual Question AnsweringQuestion AnsweringVisual Question Answering

Multimodal Confidence Modeling in Audio-Visual Quality Assessment

2026-05-02 · Mayesha Maliha R. Mithila, Mylene C. Q. Farias arxiv

Audio-visual quality assessment (AVQA) is essential for streaming, teleconferencing, and immersive media. In realistic streaming scenarios, distortions are often asymmetric, where one modality may be severely degraded wh…