paper-with-me

Papers

Telling Left from Right: Learning Spatial Correspondence of Sight and Sound

2020-06-11 · CVPR 2020 6 · Karren Yang, Bryan Russell, Justin Salamon

Self-supervised audio-visual learning aims to capture useful representations of video by leveraging correspondences between visual and audio inputs. Existing approaches have focused primarily on matching semantic information between the sensory streams. We propose a novel self-supervised task to leverage an orthogonal principle: matching spatial information in the audio stream to the positions of sound sources in the visual stream. Our approach is simple yet effective. We train a model to determine whether the left and right audio channels have been flipped, forcing it to reason about spatial localization across the visual and audio streams. To train and evaluate our method, we introduce a large-scale video dataset, YouTube-ASMR-300K, with spatial audio comprising over 900 hours of footage. We demonstrate that understanding spatial correspondence enables models to perform better on three audio-visual tasks, achieving quantitative gains over supervised and self-supervised baselines that do not leverage spatial audio cues. We also show how to extend our self-supervised approach to 360 degree videos with ambisonic audio.

📄 PDF Abstract BibTeX arXiv:2006.06175

Code (0)

등록된 구현이 없습니다.

Tasks

audio-visual learning

Similar Papers 제목 키워드 기반

Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence

2023-11-28 · CVPR 2024 1 · Junyi Zhang, Charles Herrmann, Junhwa Hur, Eric Chen 외

While pre-trained large-scale vision models have shown significant promise for semantic correspondence, their features often struggle to grasp the geometry and orientation of instances. This paper identifies the importan…

Animal Pose EstimationPose EstimationSemantic correspondence

Left-Right Symmetry Breaking in CLIP-style Vision-Language Models Trained on Synthetic Spatial-Relation Data

2026-01-19 · Takaki Yamamoto, Chihiro Noguchi, Toshihiro Tanizawa arxiv

Spatial understanding remains a key challenge in vision-language models. Yet it is still unclear whether such understanding is truly acquired, and if so, through what mechanisms. We present a controllable 1D image-text t…

Self-supervised Audio Spatialization with Correspondence Classifier

2019-05-14 · Yu-Ding Lu, Hsin-Ying Lee, Hung-Yu Tseng, Ming-Hsuan Yang

Spatial audio is an essential medium to audiences for 3D visual and auditory experience. However, the recording devices and techniques are expensive or inaccessible to the general public. In this work, we propose a self-…

LeftRefill: Filling Right Canvas based on Left Reference through Generalized Text-to-Image Diffusion Model

2023-05-19 · CVPR 2024 1 · Chenjie Cao, Yunuo Cai, Qiaole Dong, Yikai Wang 외

This paper introduces LeftRefill, an innovative approach to efficiently harness large Text-to-Image (T2I) diffusion models for reference-guided image synthesis. As the name implies, LeftRefill horizontally stitches refer…

Image GenerationImage InpaintingImage ManipulationNovel View Synthesis+1

What's left can't be right -- The remaining positional incompetence of contrastive vision-language models

2023-11-20 · Nils Hoehing, Ellen Rushe, Anthony Ventresque

Contrastive vision-language models like CLIP have been found to lack spatial understanding capabilities. In this paper we discuss the possible causes of this phenomenon by analysing both datasets and embedding space. By …