paper-with-me

Papers

AnyMatch: Supercharging Universal Multi-Modal Image Matching with Large-Scale Single-View Images

2026-06-30 · Meng Yang, Zizhuo Li, Linfeng Tang, Fan Fan, Jiayi Ma arxiv

Multi-modal image matching is essential for visual localization and multi-sensor fusion, but it is hindered by the scarcity of large-scale training data with precise geometric annotations. Existing real-world datasets suffer from prohibitive costs, limited scene diversity, and errors in SfM-MVS pipelines, while synthetic methods struggle to maintain 3D geometric consistency or achieve photorealistic appearance. To address this, we propose AnyMatch, a novel framework that leverages abundant, easily accessible single-view images at minimal cost to generate rich multi-modal training data. AnyMatch integrates monocular depth estimation, 3D reprojection, diffusion-based inpainting, and crossmodal image translation to synthesize multi-view, multi-modal image pairs with 3D geometric fidelity. Crucially, our method provides annotations that strictly adhere to 3D geometric consistency through explicit 3D reprojection, avoiding SfM-MVS error accumulation. Furthermore, AnyMatch offers strong scalability, enabling controllable scene diversity and annotation difficulty via adjustable input and camera parameters. We construct Any-syn, a large-scale synthetic multi-modal dataset using AnyMatch. Experimental results show that matching networks (e.g., LoFTR, EDM, RoMa) fine-tuned on Any-syn achieve substantial performance gains on multi-modal benchmarks, exhibiting superior generalization and robustness compared to models trained on existing data.

📄 PDF Abstract BibTeX arXiv:2606.31077

Code (0)

등록된 구현이 없습니다.

Tasks

Monocular Depth EstimationVisual LocalizationImage Matching

Similar Papers 제목 키워드 기반

AnyMatch -- Efficient Zero-Shot Entity Matching with a Small Language Model

2024-09-06 · Zeyu Zhang, Paul Groth, Iacer Calixto, Sebastian Schelter

Entity matching (EM) is the problem of determining whether two records refer to same real-world entity, which is crucial in data integration, e.g., for product catalogs or address databases. A major drawback of many EM a…

AttributeAutoMLData IntegrationLanguage Modeling+3

AutoGluon-Multimodal (AutoMM): Supercharging Multimodal AutoML with Foundation Models

2024-04-24 · Zhiqiang Tang, Haoyang Fang, Su Zhou, Taojiannan Yang 외

AutoGluon-Multimodal (AutoMM) is introduced as an open-source AutoML library designed specifically for multimodal learning. Distinguished by its exceptional ease of use, AutoMM enables fine-tuning of foundation models wi…

AutoMLImage Segmentationobject-detectionObject Detection+2

Supercharging Thermal Gaussian Splatting with Depth Estimation

2026-05-28 · Manoj Biswanath, Chenxin Cai, Hannah Schieber, Daniel Roth 외 arxiv

Efficient and robust 3D scene representation is crucial in autonomous driving, robotics, and related fields. While RGB images provide valuable content for 3D reconstruction, other modalities like thermal or depth can ena…

Novel View SynthesisAutonomous Driving3D ReconstructionDepth Estimation

Euclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions

2024-12-11 · Jiarui Zhang, Ollie Liu, Tianyu Yu, Jinyi Hu 외

Multimodal large language models (MLLMs) have made rapid progress in recent years, yet continue to struggle with low-level visual perception (LLVP) -- particularly the ability to accurately describe the geometric details…

Medical Image Analysis

Seeing and Reasoning with Confidence: Supercharging Multimodal LLMs with an Uncertainty-Aware Agentic Framework

2025-03-11 · Zhuo Zhi, Chen Feng, Adam Daneshmend, Mine Orlu 외

Multimodal large language models (MLLMs) show promise in tasks like visual question answering (VQA) but still face challenges in multimodal reasoning. Recent works adapt agentic frameworks or chain-of-thought (CoT) reaso…

Conformal PredictionMultimodal ReasoningQuestion AnsweringUncertainty Quantification+2