paper-with-me

Papers

Structure-to-Image: Zero-Shot Depth Estimation in Colonoscopy via High-Fidelity Sim-to-Real Adaptation

2026-02-25 · Juan Yang, Yuyan Zhang, Han Jia, Bing Hu, Wanzhong Song arxiv

Monocular depth estimation (MDE) for colonoscopy is hampered by the domain gap between simulated and real-world images. Existing image-to-image translation methods, which use depth as a posterior constraint, often produce structural distortions and specular highlights by failing to balance realism with structure consistency. To address this, we propose a Structure-to-Image paradigm that transforms the depth map from a passive constraint into an active generative foundation. We are the first to introduce phase congruency to colonoscopic domain adaptation and design a cross-level structure constraint to co-optimize geometric structures and fine-grained details like vascular textures. In zero-shot evaluations conducted on a publicly available phantom dataset, the MDE model that was fine-tuned on our generated data achieved a maximum reduction of 44.18% in RMSE compared to competing methods. Our code is available at https://github.com/YyangJJuan/PC-S2I.git.

📄 PDF Abstract BibTeX arXiv:2602.21740

Code (0)

등록된 구현이 없습니다.

Tasks

Monocular Depth EstimationImage-to-Image TranslationDomain Adaptation

Similar Papers 제목 키워드 기반

Metric3Dv2: A Versatile Monocular Geometric Foundation Model for Zero-shot Metric Depth and Surface Normal Estimation

2024-03-22 · Under review for Transaction 2024 4 · Mu Hu, Wei Yin, Chi Zhang, Zhipeng Cai 외

We introduce Metric3D v2, a geometric foundation model for zero-shot metric depth and surface normal estimation from a single image, which is crucial for metric 3D recovery. While depth and normal are geometrically relat…

Depth EstimationSurface Normal EstimationZero-shot Generalization

Can Language Understand Depth?

2022-07-03 · Renrui Zhang, Ziyao Zeng, Ziyu Guo, Yafeng Li

Besides image classification, Contrastive Language-Image Pre-training (CLIP) has accomplished extraordinary success for a wide range of vision tasks, including object-level and 3D space understanding. However, it's still…

Depth Estimationimage-classificationImage ClassificationMonocular Depth Estimation

GRIN: Zero-Shot Metric Depth with Pixel-Level Diffusion

2024-09-15 · Vitor Guizilini, Pavel Tokmakov, Achal Dave, Rares Ambrus

3D reconstruction from a single image is a long-standing problem in computer vision. Learning-based methods address its inherent scale ambiguity by leveraging increasingly large labeled and unlabeled datasets, to produce…

3D ReconstructionDepth EstimationImage GenerationMonocular Depth Estimation

Metric3D: Towards Zero-shot Metric 3D Prediction from A Single Image

2023-07-20 · ICCV 2023 1 · Wei Yin, Chi Zhang, Hao Chen, Zhipeng Cai 외

Reconstructing accurate 3D scenes from images is a long-standing vision task. Due to the ill-posedness of the single-image reconstruction problem, most well-established methods are built upon multi-view geometry. State-o…

Depth EstimationImage ReconstructionMonocular Depth EstimationZero-shot Generalization

Structure-Aware Radar-Camera Depth Estimation

2025-06-05 · Fuyi Zhang, Zhu Yu, Chunhao Li, Runmin Zhang 외

Monocular depth estimation aims to determine the depth of each pixel from an RGB image captured by a monocular camera. The development of deep learning has significantly advanced this field by facilitating the learning o…

Depth EstimationMonocular Depth Estimationregression