paper-with-me

홈 › Papers

Local Area Transform for Cross-Modality Correspondence Matching and Deep Scene Recognition

2019-01-03 · Seungchul Ryu

Establishing correspondences is a fundamental task in variety of image processing and computer vision applications. In particular, finding the correspondences between a non-linearly deformed image pair induced by different modality conditions is a challenging problem. This paper describes a efficient but powerful image transform called local area transform (LAT) for modality-robust correspondence estimation. Specifically, LAT transforms an image from the intensity domain to the local area domain, which is invariant under nonlinear intensity deformations, especially radiometric, photometric, and spectral deformations. In addition, robust feature descriptors are reformulated with LAT for several practical applications. Furthermore, LAT-convolution layer and Aception block are proposed and, with these novel components, deep neural network called LAT-Net is proposed especially for scene recognition task. Experimental results show that LATransformed images provide a consistency for nonlinearly deformed images, even under random intensity deformations. LAT reduces the mean absolute difference as compared to conventional methods. Furthermore, the reformulation of descriptors with LAT shows superiority to conventional methods, which is a promising result for the tasks of cross-spectral and modality correspondence matching. the local area can be considered as an alternative domain to the intensity domain to achieve robust correspondence matching, image recognition, and a lot of applications: such as feature matching, stereo matching, dense correspondence matching, image recognition, and image retrieval.

📄 PDF Abstract BibTeX arXiv:1901.00927

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalRetrievalScene RecognitionStereo MatchingStereo Matching Hand

Similar Papers 제목 키워드 기반

SC3EF: A Joint Self-Correlation and Cross-Correspondence Estimation Framework for Visible and Thermal Image Registration

2025-04-17 · Xi Tong, Xing Luo, Jiangxin Yang, Yanpeng Cao

Multispectral imaging plays a critical role in a range of intelligent transportation applications, including advanced driver assistance systems (ADAS), traffic monitoring, and night vision. However, accurate visible and …

Image RegistrationOptical Flow Estimation

Unsupervised Visible-Infrared Person Re-Identification via Progressive Graph Matching and Alternate Learning

2023-01-01 · CVPR 2023 1 · Zesen Wu, Mang Ye

Unsupervised visible-infrared person re-identification is a challenging task due to the large modality gap and the unavailability of cross-modality correspondences. Cross-modality correspondences are very crucial to …

Contrastive LearningGraph MatchingPerson Re-Identification

UniCorrn: Unified Correspondence Transformer Across 2D and 3D

2026-05-05 · Prajnan Goswami, Tianye Ding, Feng Liu, Huaizu Jiang arxiv

Visual correspondence across image-to-image (2D-2D), image-to-point cloud (2D-3D), and point cloud-to-point cloud (3D-3D) geometric matching forms the foundation for numerous 3D vision tasks. Despite sharing a similar pr…

Geometric MatchingPoint Clouds

All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation

2026-09-23 · Ohad Rahamim, Dvir Samuel, Idan Schwartz, Gal Chechik hf

Video is a rich representation of a physical event, capturing appearance, geometry, motion, and temporal evolution. Other modalities, such as 3D body motion or audio, encode narrower aspects of the same event. We find th…

Audio GenerationVideo Generation

VoxCor: Training-Free Volumetric Features for Multimodal Voxel Correspondence

2026-05-13 · Guney Tombak, Ertunc Erdil, Ender Konukoglu arxiv

Cross-modal 3D medical image analysis requires voxelwise representations that remain anatomically consistent across imaging contrasts, scanners, and acquisition protocols. Recent work has shown that frozen 2D Vision Tran…