paper-with-me

Papers

XoFTR: Cross-modal Feature Matching Transformer

2024-04-15 · Önder Tuzcuoğlu, Aybora Köksal, Buğra Sofu, Sinan Kalkan, A. Aydin Alatan

We introduce, XoFTR, a cross-modal cross-view method for local feature matching between thermal infrared (TIR) and visible images. Unlike visible images, TIR images are less susceptible to adverse lighting and weather conditions but present difficulties in matching due to significant texture and intensity differences. Current hand-crafted and learning-based methods for visible-TIR matching fall short in handling viewpoint, scale, and texture diversities. To address this, XoFTR incorporates masked image modeling pre-training and fine-tuning with pseudo-thermal image augmentation to handle the modality differences. Additionally, we introduce a refined matching pipeline that adjusts for scale discrepancies and enhances match reliability through sub-pixel level refinement. To validate our approach, we collect a comprehensive visible-thermal dataset, and show that our method outperforms existing methods on many benchmarks.

📄 PDF Abstract BibTeX arXiv:2404.09692

Code (1)

ondert/xoftr 공식 구현 pytorch

Tasks

Image Augmentation

Similar Papers 제목 키워드 기반

Are Pretrained Image Matchers Good Enough for SAR-Optical Satellite Registration?

2026-04-11 · Isaac Corley, Alex Stoken, Gabriele Berton arxiv

Cross-modal optical-SAR (Synthetic Aperture Radar) registration is a bottleneck for disaster-response via remote sensing, yet modern image matchers are developed and benchmarked almost exclusively on natural-image domain…

Domain AdaptationImage Matching

UFM: Unified Feature Matching Pre-training with Multi-Modal Image Assistants

2025-03-26 · Yide Di, Yun Liao, Hao Zhou, Kaijun Zhu 외

Image feature matching, a foundational task in computer vision, remains challenging for multimodal image applications, often necessitating intricate training on specific datasets. In this paper, we introduce a Unified Fe…

Data Augmentation

CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language Transformers

2023-05-27 · Dachuan Shi, Chaofan Tao, Anyi Rao, Zhendong Yang 외

Recent vision-language models have achieved tremendous advances. However, their computational costs are also escalating dramatically, making model acceleration exceedingly critical. To pursue more efficient vision-langua…

Image CaptioningImage RetrievalImage-text RetrievalImage-to-Text Retrieval+5

Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey

2025-07-30 · Weide Liu, Wei Zhou, Jun Liu, Ping Hu 외 arxiv

Feature matching is a cornerstone task in computer vision, essential for applications such as image retrieval, stereo matching, 3D reconstruction, and SLAM. This survey comprehensively reviews modality-based feature matc…

Medical Image Registration3D ReconstructionImage RetrievalImage Matching

UniCorrn: Unified Correspondence Transformer Across 2D and 3D

2026-05-05 · Prajnan Goswami, Tianye Ding, Feng Liu, Huaizu Jiang arxiv

Visual correspondence across image-to-image (2D-2D), image-to-point cloud (2D-3D), and point cloud-to-point cloud (3D-3D) geometric matching forms the foundation for numerous 3D vision tasks. Despite sharing a similar pr…

Geometric MatchingPoint Clouds