paper-with-me

Papers

Audio-visual training for improved grounding in video-text LLMs

2024-07-21 · Shivprasad Sagare, Hemachandran S, Kinshuk Sarabhai, Prashant Ullegaddi, Rajeshkumar SA

Recent advances in multimodal LLMs, have led to several video-text models being proposed for critical video-related tasks. However, most of the previous works support visual input only, essentially muting the audio signal in the video. Few models that support both audio and visual input, are not explicitly trained on audio data. Hence, the effect of audio towards video understanding is largely unexplored. To this end, we propose a model architecture that handles audio-visual inputs explicitly. We train our model with both audio and visual data from a video instruction-tuning dataset. Comparison with vision-only baselines, and other audio-visual models showcase that training on audio data indeed leads to improved grounding of responses. For better evaluation of audio-visual models, we also release a human-annotated benchmark dataset, with audio-aware question-answer pairs.

📄 PDF Abstract BibTeX arXiv:2407.15046

Code (0)

등록된 구현이 없습니다.

Tasks

Video Understanding

Similar Papers 제목 키워드 기반

Target-Aware Spatio-Temporal Reasoning via Answering Questions in Dynamics Audio-Visual Scenarios

2023-05-21 · Yuanyuan Jiang, Jianqin Yin

Audio-visual question answering (AVQA) is a challenging task that requires multistep spatio-temporal reasoning over multimodal contexts. Recent works rely on elaborate target-agnostic parsing of audio-visual scenes for s…

Audio-visual Question AnsweringAudio-Visual Question Answering (AVQA)Question AnsweringScene Understanding+1

Video-Guided Curriculum Learning for Spoken Video Grounding

2022-09-01 · Yan Xia, Zhou Zhao, Shangwei Ye, Yang Zhao 외

In this paper, we introduce a new task, spoken video grounding (SVG), which aims to localize the desired video fragments from spoken language descriptions. Compared with using text, employing audio requires the model to …

Video Grounding

Cross-Modal learning for Audio-Visual Video Parsing

2021-04-03 · Jatin Lamba, abhishek, Jayaprakash Akula, Rishabh Dabral 외

In this paper, we present a novel approach to the audio-visual video parsing (AVVP) task that demarcates events from a video separately for audio and visual modalities. The proposed parsing approach simultaneously detect…

Event DetectionMultiple Instance LearningVideo Grounding

ChronusOmni: Improving Time Awareness of Omni Large Language Models

2025-12-10 · Yijing Chen, Yihan Wu, Kaisi Guan, Yuchen Ren 외 arxiv

Time awareness is a fundamental ability of omni large language models, especially for understanding long videos and answering complex questions. Previous approaches mainly target vision-language scenarios and focus on th…

Reinforcement Learning

Pano-AVQA: Grounded Audio-Visual Question Answering on 360deg Videos

2021-01-01 · ICCV 2021 10 · Heeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee 외

360deg videos convey holistic views for the surroundings of a scene. It provides audio-visual cues beyond predetermined normal field of views and displays distinctive spatial relations on a sphere. However, previous …

Audio-visual Question AnsweringQuestion AnsweringRelationVisual Question Answering+1