paper-with-me

Papers

Self-Supervised Spatial Correspondence Across Modalities

2025-01-01 · CVPR 2025 1 · Ayush Shrivastava, Andrew Owens

We present a method for finding cross-modal space-time correspondences. Given two images from different visual modalities, such as an RGB image and a depth map, our model identifies which pairs of pixels correspond to the same physical points in the scene. To solve this problem, we extend the contrastive random walk framework to simultaneously learn cycle-consistent feature representations for both cross-modal and intra-modal matching. The resulting model is simple and has no explicit photo-consistency assumptions. It can be trained entirely using unlabeled data, without the need for any spatially aligned multimodal image pairs. We evaluate our method on both geometric and semantic correspondence tasks. For geometric matching, we consider challenging tasks such as RGB-to-depth and RGB-to-thermal matching (and vice versa); for semantic matching, we evaluate on photo-sketch and cross-style image alignment. Our method achieves strong performance across all benchmarks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Geometric MatchingSemantic correspondence

Similar Papers 제목 키워드 기반

PointCMC: Cross-Modal Multi-Scale Correspondences Learning for Point Cloud Understanding

2022-11-22 · Honggu Zhou, Xiaogang Peng, Jiawei Mao, Zizhao Wu 외

Some self-supervised cross-modal learning approaches have recently demonstrated the potential of image signals for enhancing point cloud representation. However, it remains a question on how to directly model cross-modal…

3D Object ClassificationRepresentation Learning

Beyond Instance-Level Self-Supervision in 3D Multi-Modal Medical Imaging

2026-05-14 · Tan Pan, Shuhao Mei, Yixuan Sun, Kaiyu Guo 외 arxiv

Self-supervised pre-training methods in medical imaging typically treat each individual as an isolated instance, learning representations through augmentation-based objectives or masked reconstruction. They often do not …

Telling Left from Right: Learning Spatial Correspondence of Sight and Sound

2020-06-11 · CVPR 2020 6 · Karren Yang, Bryan Russell, Justin Salamon

Self-supervised audio-visual learning aims to capture useful representations of video by leveraging correspondences between visual and audio inputs. Existing approaches have focused primarily on matching semantic informa…

audio-visual learning

Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos

2023-07-10 · CVPR 2024 1 · Sagnik Majumder, Ziad Al-Halah, Kristen Grauman

We propose a self-supervised method for learning representations based on spatial audio-visual correspondences in egocentric videos. Our method uses a masked auto-encoding framework to synthesize masked binaural (multi-c…

Active Speaker DetectionAudio DenoisingDenoising

Correspondence-Oriented Imitation Learning: Flexible Visuomotor Control with 3D Conditioning

2025-12-05 · Yunhao Cao, Zubin Bhaumik, Jessie Jia, Xingyi He 외 arxiv

We introduce Correspondence-Oriented Imitation Learning (COIL), a conditional policy learning framework for visuomotor control with a flexible task representation in 3D. At the core of our approach, each task is defined …