paper-with-me

Papers

ColonAdapter: Geometry Estimation Through Foundation Model Adaptation for Colonoscopy

2025-11-27 · Zhiyi Jiang, Yifu Wang, Xuelian Cheng, Zongyuan Ge arxiv

Estimating 3D geometry from monocular colonoscopy images is challenging due to non-Lambertian surfaces, moving light sources, and large textureless regions. While recent 3D geometric foundation models eliminate the need for multi-stage pipelines, their performance deteriorates in clinical scenes. These models are primarily trained on natural scene datasets and struggle with specularity and homogeneous textures typical in colonoscopy, leading to inaccurate geometry estimation. In this paper, we present ColonAdapter, a self-supervised fine-tuning framework that adapts geometric foundation models for colonoscopy geometry estimation. Our method leverages pretrained geometric priors while tailoring them to clinical data. To improve performance in low-texture regions and ensure scale consistency, we introduce a Detail Restoration Module (DRM) and a geometry consistency loss. Furthermore, a confidence-weighted photometric loss enhances training stability in clinical environments. Experiments on both synthetic and real datasets demonstrate that our approach achieves state-of-the-art performance in camera pose estimation, monocular depth prediction, and dense 3D point map reconstruction, without requiring ground-truth intrinsic parameters.

📄 PDF Abstract BibTeX arXiv:2511.22250

Code (0)

등록된 구현이 없습니다.

Tasks

Camera Pose Estimation

Similar Papers 제목 키워드 기반

Geometry-Consistent Endoscopic Representations for Image-Guided Navigation via Structured Foundation Model Adaptation

2026-06-15 · Hongchao Shu, Roger D. Soberanis-Mukul, Hao Ding, Morgan Ringel 외 arxiv

Accurate vision-based navigation in monocular endoscopy is difficult due to limited depth cues, weak tissue texture, non-rigid deformation, and substantial appearance variation across domains, all of which complicate pos…

Monocular Depth EstimationRepresentation LearningPose Estimation

Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models

2026-08-11 · Seokhyun Youn, Dahyeon Kye, Sung-Ho Bae, Jihyong Oh hf

Recent Vision Foundation Models (VFMs) predict depth, camera pose, and pointmap in a single forward pass without per-scene optimization, achieving strong generalization. However, enforcing explicit multi-view geometric c…

Test-time Adaptation

Emergent Extreme-View Geometry in 3D Foundation Models

2025-11-27 · Yiwen Zhang, Joseph Tung, Ruojin Cai, David Fouhey 외 arxiv

3D foundation models (3DFMs) have recently transformed 3D vision, enabling joint prediction of depths, poses, and point maps directly from images. Yet their ability to reason under extreme, non-overlapping views remains …

3D ReconstructionPose Estimation

DrivingDepth: Sparse-Prompted Pixel-wise Scale Correction for Driving Depth Estimation

2026-06-30 · Chi Huang, Wenhao Zhang, Hang Yin, YuAn Wang 외 arxiv

Dense depth estimation for autonomous driving faces a geometry-scale conflict: depth foundation models deliver pixel-aligned dense visual geometry without reliable metric scale, while projected LiDAR provides metric anch…

Autonomous DrivingDepth Estimation

Endo-FASt3r: Endoscopic Foundation model Adaptation for Structure from motion

2025-03-10 · Mona Sheikh Zeinoddin, Mobarakol Islam, Zafer Tandogdu, Greg Shaw 외

Accurate depth and camera pose estimation is essential for achieving high-quality 3D visualisations in robotic-assisted surgery. Despite recent advancements in foundation model adaptation to monocular depth estimation of…

Camera Pose EstimationDepth EstimationMonocular Depth EstimationPose Estimation+1