paper-with-me

Papers

Past and Future Motion Guided Network for Audio Visual Event Localization

2022-05-08 · Tingxiu Chen, Jianqin Yin, Jin Tang

In recent years, audio-visual event localization has attracted much attention. It's purpose is to detect the segment containing audio-visual events and recognize the event category from untrimmed videos. Existing methods use audio-guided visual attention to lead the model pay attention to the spatial area of the ongoing event, devoting to the correlation between audio and visual information but ignoring the correlation between audio and spatial motion. We propose a past and future motion extraction (pf-ME) module to mine the visual motion from videos ,embedded into the past and future motion guided network (PFAGN), and motion guided audio attention (MGAA) module to achieve focusing on the information related to interesting events in audio modality through the past and future visual motion. We choose AVE as the experimental verification dataset and the experiments show that our method outperforms the state-of-the-arts in both supervised and weakly-supervised settings.

📄 PDF Abstract BibTeX arXiv:2205.03802

Code (0)

등록된 구현이 없습니다.

Tasks

audio-visual event localization

Similar Papers 제목 키워드 기반

Visually Guided Self Supervised Learning of Speech Representations

2020-01-13 · Abhinav Shukla, Konstantinos Vougioukas, Pingchuan Ma, Stavros Petridis 외

Self supervised representation learning has recently attracted a lot of research interest for both the audio and visual modalities. However, most works typically focus on a particular modality or feature alone and there …

Emotion RecognitionRepresentation LearningSelf-Supervised LearningSpeech Emotion Recognition+2

Sound2Sight: Generating Visual Dynamics from Sound and Context

2020-07-23 · ECCV 2020 8 · Anoop Cherian, Moitreya Chatterjee, Narendra Ahuja

Learning associations across modalities is critical for robust multimodal reasoning, especially when a modality may be missing during inference. In this paper, we study this problem in the context of audio-conditioned vi…

Multimodal ReasoningVideo Forecasting

MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation

2024-12-05 · Longtao Zheng, Yifan Zhang, Hanzhong Guo, Jiachun Pan 외

Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, maintaining long-term identity consistency…

Portrait AnimationVideo Generation

Past Movements-Guided Motion Representation Learning for Human Motion Prediction

2024-08-04 · Junyu Shi, Baoxuan Wang

Human motion prediction based on 3D skeleton is a significant challenge in computer vision, primarily focusing on the effective representation of motion. In this paper, we propose a self-supervised learning framework des…

Human motion predictionmotion predictionPredictionRepresentation Learning+1

Learning to Hear by Seeing: It's Time for Vision Language Models to Understand Artistic Emotion from Sight and Sound

2025-11-15 · Dengming Zhang, Weitao You, Jingxiong Li, Weishen Lin 외 arxiv

Emotion understanding is critical for making Large Language Models (LLMs) more general, reliable, and aligned with humans. Art conveys emotion through the joint design of visual and auditory elements, yet most prior work…