paper-with-me

Papers

Pano-AVQA: Grounded Audio-Visual Question Answering on 360$^\circ$ Videos

2021-10-11 · Heeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee, Gunhee Kim

360$^\circ$ videos convey holistic views for the surroundings of a scene. It provides audio-visual cues beyond pre-determined normal field of views and displays distinctive spatial relations on a sphere. However, previous benchmark tasks for panoramic videos are still limited to evaluate the semantic understanding of audio-visual relationships or spherical spatial property in surroundings. We propose a novel benchmark named Pano-AVQA as a large-scale grounded audio-visual question answering dataset on panoramic videos. Using 5.4K 360$^\circ$ video clips harvested online, we collect two types of novel question-answer pairs with bounding-box grounding: spherical spatial relation QAs and audio-visual relation QAs. We train several transformer-based models from Pano-AVQA, where the results suggest that our proposed spherical spatial embeddings and multimodal training objectives fairly contribute to a better semantic understanding of the panoramic surroundings on the dataset.

📄 PDF Abstract BibTeX arXiv:2110.05122

Code (1)

hs-yn/panoavqa 공식 구현 pytorch

Tasks

Audio-visual Question AnsweringQuestion AnsweringRelationVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Pano-AVQA: Grounded Audio-Visual Question Answering on 360deg Videos

2021-01-01 · ICCV 2021 10 · Heeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee 외

360deg videos convey holistic views for the surroundings of a scene. It provides audio-visual cues beyond predetermined normal field of views and displays distinctive spatial relations on a sphere. However, previous …

Audio-visual Question AnsweringQuestion AnsweringRelationVisual Question Answering+1

Learning to Answer Questions in Dynamic Audio-Visual Scenarios

2022-03-26 · CVPR 2022 1 · Guangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu 외

In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in videos. The problem requires comprehensive …

audio-visual learningAudio-visual Question AnsweringAudio-Visual Question Answering (AVQA)AUDIO-VISUAL QUESTION ANSWERING (MUSIC-AVQA-v2.0)+4

Multi-Modal Scene Graph with Kolmogorov-Arnold Experts for Audio-Visual Question Answering

2025-11-28 · Zijian Fu, Changsheng Lv, Mengshi Qi, Huadong Ma arxiv

In this paper, we propose a novel Multi-Modal Scene Graph with Kolmogorov-Arnold Expert Network for Audio-Visual Question Answering (SHRIKE). The task aims to mimic human reasoning by extracting and fusing information fr…

Audio-visual Question Answering

SaSR-Net: Source-Aware Semantic Representation Network for Enhancing Audio-Visual Question Answering

2024-11-07 · Tianyu Yang, Yiyang Nan, Lisen Dai, Zhenwen Liang 외

Audio-Visual Question Answering (AVQA) is a challenging task that involves answering questions based on both auditory and visual information in videos. A significant challenge is interpreting complex multi-modal scenes, …

Audio-visual Question AnsweringAudio-Visual Question Answering (AVQA)Question AnsweringVisual Question Answering

AVQACL: A Novel Benchmark for Audio-Visual Question Answering Continual Learning

2025-01-01 · CVPR 2025 1 · Kaixuan Wu, Xinde Li, Xinling Li, Chuanfei Hu 외

In this paper, a novel benchmark for audio-visual question answering continual learning (AVQACL) is introduced, aiming to study fine-grained scene understanding and spatial-temporal reasoning in videos under a contin…

Audio-visual Question AnsweringContinual LearningKnowledge DistillationQuestion Answering+2