paper-with-me

Papers

Enhancing Generalization of Depth Estimation Foundation Model via Weakly-Supervised Adaptation with Regularization

2025-11-18 · Yan Huang, Yongyi Su, Xin Lin, Le Zhang, Xun Xu arxiv

The emergence of foundation models has substantially advanced zero-shot generalization in monocular depth estimation (MDE), as exemplified by the Depth Anything series. However, given access to some data from downstream tasks, a natural question arises: can the performance of these models be further improved? To this end, we propose WeSTAR, a parameter-efficient framework that performs Weakly supervised Self-Training Adaptation with Regularization, designed to enhance the robustness of MDE foundation models in unseen and diverse domains. We first adopt a dense self-training objective as the primary source of structural self-supervision. To further improve robustness, we introduce semantically-aware hierarchical normalization, which exploits instance-level segmentation maps to perform more stable and multi-scale structural normalization. Beyond dense supervision, we introduce a cost-efficient weak supervision in the form of pairwise ordinal depth annotations to further guide the adaptation process, which enforces informative ordinal constraints to mitigate local topological errors. Finally, a weight regularization loss is employed to anchor the LoRA updates, ensuring training stability and preserving the model's generalizable knowledge. Extensive experiments on both realistic and corrupted out-of-distribution datasets under diverse and challenging scenarios demonstrate that WeSTAR consistently improves generalization and achieves state-of-the-art performance across a wide range of benchmarks.

📄 PDF Abstract BibTeX arXiv:2511.14238

Code (0)

등록된 구현이 없습니다.

Tasks

Monocular Depth EstimationZero-shot Generalization

Similar Papers 제목 키워드 기반

Weakly-supervised Pre-training for 3D Human Pose Estimation via Perspective Knowledge

2022-11-22 · Zhongwei Qiu, Kai Qiu, Jianlong Fu, Dongmei Fu

Modern deep learning-based 3D pose estimation approaches require plenty of 3D pose annotations. However, existing 3D datasets lack diversity, which limits the performance of current methods and their generalization abili…

3D Human Pose Estimation3D Pose EstimationDiversityPose Estimation

EndoUFM: Utilizing Foundation Models for Monocular depth estimation of endoscopic images

2025-08-25 · Xinning Yao, Bo Liu, Bojian Li, Jingjing Wang 외 arxiv

Depth estimation is a foundational component for 3D reconstruction in minimally invasive endoscopic surgeries. However, existing monocular depth estimation techniques often exhibit limited performance to the varying illu…

Monocular Depth Estimation3D Reconstruction

Towards Depth Foundation Model: Recent Trends in Vision-Based Depth Estimation

2025-07-15 · Zhen Xu, HongYu Zhou, Sida Peng, Haotong Lin 외

Depth estimation is a fundamental task in 3D computer vision, crucial for applications such as 3D reconstruction, free-viewpoint rendering, robotics, autonomous driving, and AR/VR technologies. Traditional methods relyin…

3D ReconstructionAutonomous DrivingDepth EstimationZero-shot Generalization

Boosting Weakly Supervised Object Detection using Fusion and Priors from Hallucinated Depth

2023-03-20 · Cagri Gungor, Adriana Kovashka

Despite recent attention and exploration of depth for various tasks, it is still an unexplored modality for weakly-supervised object detection (WSOD). We propose an amplifier method for enhancing the performance of WSOD …

Depth EstimationMonocular Depth EstimationMultiple Instance Learningobject-detection+2

DEFOM-Stereo: Depth Foundation Model Based Stereo Matching

2025-01-16 · CVPR 2025 1 · Hualie Jiang, Zhiqiang Lou, Laiyan Ding, Rui Xu 외

Stereo matching is a key technique for metric depth estimation in computer vision and robotics. Real-world challenges like occlusion and non-texture hinder accurate disparity estimation from binocular matching cues. Rece…

Depth EstimationDisparity EstimationmodelStereo Matching+1