audio-visual event localization
1개 벤치마크 · 논문 32편 · 이 태스크의 논문 보기 →
Benchmarks
UnAV-100
Most implemented
Positive Sample Propagation along the Audio-Visual Event Line
Dual-modality seq2seq network for audio-visual event localization
Audio-Visual Event Localization in Unconstrained Videos
Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration
Towards Open-Vocabulary Audio-Visual Event Localization
Papers
Hierarchical Semantic-Constrained Heterogeneous Graph for Audio-Visual Event Localization
Open-vocabulary audio-visual event localization (OV-AVEL) jointly models audio-visual cues to recognize and temporally localize events, including categories unseen during training. Existing methods primarily learn joint …
audio-visual event localizationRA-SSU: Towards Fine-Grained Audio-Visual Learning with Region-Aware Sound Source Understanding
Audio-Visual Learning (AVL) is one fundamental task of multi-modality learning and embodied intelligence, displaying the vital role in scene understanding and interaction. However, previous researchers mostly focus on ex…
audio-visual event localizationSound Source LocalizationScene UnderstandingMoLT: Mixture of Layer-Wise Tokens for Efficient Audio-Visual Learning
In this paper, we propose Mixture of Layer-Wise Tokens (MoLT), a parameter- and memory-efficient adaptation framework for audio-visual learning. The key idea of MoLT is to replace conventional, computationally heavy sequ…
audio-visual event localizationAudio-visual Question AnsweringReal-Time Inference for Distributed Multimodal Systems under Communication Delay Uncertainty
Connected cyber-physical systems perform inference based on real-time inputs from multiple data streams. Uncertain communication delays across data streams challenge the temporal flow of the inference process. State-of-t…
audio-visual event localizationCLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
The Dense Audio-Visual Event Localization (DAVEL) task aims to temporally localize events in untrimmed videos that occur simultaneously in both the audio and visual modalities. This paper explores DAVEL under a new and m…
audio-visual event localizationESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
Dense audio-visual event localization (DAVE) aims to identify event categories and locate the temporal boundaries in untrimmed videos. Most studies only employ event-related semantic constraints on the final outputs, lac…
audio-visual event localization