paper-with-me

Papers

Learning Representations from Audio-Visual Spatial Alignment

2020-11-03 · NeurIPS 2020 12 · Pedro Morgado, Yi Li, Nuno Vasconcelos

We introduce a novel self-supervised pretext task for learning representations from audio-visual content. Prior work on audio-visual representation learning leverages correspondences at the video level. Approaches based on audio-visual correspondence (AVC) predict whether audio and video clips originate from the same or different video instances. Audio-visual temporal synchronization (AVTS) further discriminates negative pairs originated from the same video instance but at different moments in time. While these approaches learn high-quality representations for downstream tasks such as action recognition, their training objectives disregard spatial cues naturally occurring in audio and visual signals. To learn from these spatial cues, we tasked a network to perform contrastive audio-visual spatial alignment of 360{\deg} video and spatial audio. The ability to perform spatial alignment is enhanced by reasoning over the full spatial content of the 360{\deg} video using a transformer architecture to combine representations from multiple viewpoints. The advantages of the proposed pretext task are demonstrated on a variety of audio and visual downstream tasks, including audio-visual correspondence, spatial alignment, action recognition, and video semantic segmentation.

📄 PDF Abstract BibTeX arXiv:2011.01819

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionRepresentation LearningSemantic SegmentationVideo Semantic Segmentation

Similar Papers 제목 키워드 기반

Self-supervised Learning of Audio Representations from Audio-Visual Data using Spatial Alignment

2022-06-02 · Shanshan Wang, Archontis Politis, Annamaria Mesaros, Tuomas Virtanen

Learning from audio-visual data offers many possibilities to express correspondence between the audio and visual content, similar to the human perception that relates aural and visual information. In this work, we presen…

Acoustic Scene ClassificationAction Recognitionobject-detectionObject Detection+4

Towards Modeling the Interaction of Spatial-Associative Neural Network Representations for Multisensory Perception

2018-07-13 · German I. Parisi, Jonathan Tong, Pablo Barros, Brigitte Röder 외

Our daily perceptual experience is driven by different neural mechanisms that yield multisensory interaction as the interplay between exogenous stimuli and endogenous expectations. While the interaction of multisensory c…

Causal Inference

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment

2025-05-02 · CVPR 2025 1 · Edson Araujo, Andrew Rouditchenko, Yuan Gong, Saurabhchand Bhati 외

Recent advances in audio-visual learning have shown promising results in learning representations across modalities. However, most approaches rely on global audio representations that fail to capture fine-grained tempora…

audio-visual learningcross-modal alignment

Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment

2025-03-17 · CVPR 2025 1 · Chen Liu, Peike Li, Liying Yang, Dadong Wang 외

Accurately localizing audible objects based on audio-visual cues is the core objective of audio-visual segmentation. Most previous methods emphasize spatial or temporal multi-modal modeling, yet overlook challenges from …

Contrastive Learning

SpatialV2A: Visual-Guided High-fidelity Spatial Audio Generation

2026-01-21 · Yanan Wang, Linjie Ren, Zihao Li, Junyi Wang 외 arxiv

While video-to-audio generation has achieved remarkable progress in semantic and temporal alignment, most existing studies focus solely on these aspects, paying limited attention to the spatial perception and immersive q…

Audio Generation