paper-with-me

audio-visual event localization

1개 벤치마크 · 논문 32편 · 이 태스크의 논문 보기 →

Benchmarks

UnAV-100

결과 2개

Most implemented

Papers

Hierarchical Semantic-Constrained Heterogeneous Graph for Audio-Visual Event Localization

2026-06-05 · Zhe Yang, Ruyi Zhang, Hongtao Chen, Wenrui Li 외 arxiv

Open-vocabulary audio-visual event localization (OV-AVEL) jointly models audio-visual cues to recognize and temporally localize events, including categories unseen during training. Existing methods primarily learn joint …

audio-visual event localization

RA-SSU: Towards Fine-Grained Audio-Visual Learning with Region-Aware Sound Source Understanding

2026-03-10 · Muyi Sun, Yixuan Wang, Hong Wang, Chen Su 외 arxiv

Audio-Visual Learning (AVL) is one fundamental task of multi-modality learning and embodied intelligence, displaying the vital role in scene understanding and interaction. However, previous researchers mostly focus on ex…

audio-visual event localizationSound Source LocalizationScene Understanding

MoLT: Mixture of Layer-Wise Tokens for Efficient Audio-Visual Learning

2025-11-27 · Kyeongha Rho, Hyeongkeun Lee, Jae Won Cho, Joon Son Chung arxiv

In this paper, we propose Mixture of Layer-Wise Tokens (MoLT), a parameter- and memory-efficient adaptation framework for audio-visual learning. The key idea of MoLT is to replace conventional, computationally heavy sequ…

audio-visual event localizationAudio-visual Question Answering

Real-Time Inference for Distributed Multimodal Systems under Communication Delay Uncertainty

2025-11-20 · Victor Croisfelt, João Henrique Inacio de Souza, Shashi Raj Pandey, Beatriz Soret 외 arxiv

Connected cyber-physical systems perform inference based on real-time inputs from multiple data streams. Uncertain communication delays across data streams challenge the temporal flow of the inference process. State-of-t…

audio-visual event localization

CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization

2025-08-06 · Jinxing Zhou, Ziheng Zhou, Yanghao Zhou, Yuxin Mao 외 arxiv

The Dense Audio-Visual Event Localization (DAVEL) task aims to temporally localize events in untrimmed videos that occur simultaneously in both the audio and visual modalities. This paper explores DAVEL under a new and m…

audio-visual event localization

ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization

2025-07-14 · Huilai Li, Yonghao Dang, Ying Xing, Yiming Wang 외 arxiv

Dense audio-visual event localization (DAVE) aims to identify event categories and locate the temporal boundaries in untrimmed videos. Most studies only employ event-related semantic constraints on the final outputs, lac…

audio-visual event localization

전체 32편 보기 →