paper-with-me

홈 › Papers

C3Po: Cross-View Cross-Modality Correspondence by Pointmap Prediction

2025-11-23 · Kuan Wei Huang, Brandon Li, Bharath Hariharan, Noah Snavely arxiv

Geometric models like DUSt3R have shown great advances in understanding the geometry of a scene from pairs of photos. However, they fail when the inputs are from vastly different viewpoints (e.g., aerial vs. ground) or modalities (e.g., photos vs. abstract drawings) compared to what was observed during training. This paper addresses a challenging version of this problem: predicting correspondences between ground-level photos and floor plans. Current datasets for joint photo-floor plan reasoning are limited, either lacking in varying modalities (VIGOR) or lacking in correspondences (WAFFLE). To address these limitations, we introduce a new dataset, C3, created by first reconstructing a number of scenes in 3D from Internet photo collections via structure-from-motion, then manually registering the reconstructions to floor plans gathered from the Internet, from which we can derive correspondences between images and floor plans. C3 contains 90K paired floor plans and photos across 597 scenes with 153M pixel-level correspondences and 85K camera poses. We find that state-of-the-art correspondence models struggle on this task. By training on our new data, we can improve on the best performing method by 34% in RMSE. We also use the predicted correspondences to estimate camera poses and evaluate performance using recall metrics. Lastly, we identify open challenges in cross-modal geometric reasoning that our dataset aims to help address.

📄 PDF Abstract BibTeX arXiv:2511.18559

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MV-SAM: Multi-view Promptable Segmentation using Pointmap Guidance

2026-01-25 · Yoonwoo Jeong, Cheng Sun, Yu-Chiang Frank Wang, Minsu Cho 외 arxiv

Promptable segmentation has emerged as a powerful paradigm in computer vision, enabling users to guide models in parsing complex scenes with prompts such as clicks, boxes, or textual cues. Recent advances, exemplified by…

Pointmap-Conditioned Diffusion for Consistent Novel View Synthesis

2025-01-06 · Thang-Anh-Quan Nguyen, Nathan Piasco, Luis Roldão, Moussab Bennehar 외

In this paper, we present PointmapDiffusion, a novel framework for single-image novel view synthesis (NVS) that utilizes pre-trained 2D diffusion models. Our method is the first to leverage pointmaps (i.e. rasterized 3D …

Novel View Synthesis

Pointmap Association and Piecewise-Plane Constraint for Consistent and Compact 3D Gaussian Segmentation Field

2025-02-22 · Wenhao Hu, Wenhao Chai, Shengyu Hao, Xiaotong Cui 외

Achieving a consistent and compact 3D segmentation field is crucial for maintaining semantic coherence across views and accurately representing scene structures. Previous 3D scene segmentation methods rely on video segme…

2D Panoptic Segmentation3D Scene ReconstructionPanoptic SegmentationScene Segmentation+3

C4D: 4D Made from 3D through Dual Correspondences

2025-10-16 · Shizun Wang, Zhenxiang Jiang, Xingyi Yang, Xinchao Wang arxiv

Recovering 4D from monocular video, which jointly estimates dynamic geometry and camera poses, is an inevitably challenging problem. While recent pointmap-based 3D reconstruction methods (e.g., DUSt3R) have made great pr…

Camera Pose Estimation3D ReconstructionDepth EstimationPoint Tracking

D^2USt3R: Enhancing 3D Reconstruction with 4D Pointmaps for Dynamic Scenes

2025-04-08 · Jisang Han, Honggyu An, Jaewoo Jung, Takuya Narihira 외

We address the task of 3D reconstruction in dynamic scenes, where object motions degrade the quality of previous 3D pointmap regression methods, such as DUSt3R, originally designed for static 3D scene reconstruction. Alt…

3D Reconstruction3D Scene Reconstruction