paper-with-me

홈 › Papers

DepthAgent: Towards Better Universal Depth Estimation via Sample-wise Expert Selection

2026-05-22 · Jie Zhu, Girish Chandar Ganesan, Xiaoming Liu arxiv

Monocular metric depth estimation has achieved strong progress with large-scale training and universal-camera modeling, yet robust deployment across diverse camera settings, such as perspective, fisheye, and panoramic images, remains challenging. Existing methods typically rely on a single depth estimator, overlooking that different models encode different camera assumptions and perform best under different input domains. In this paper, we show that depth experts exhibit strong sample-wise complementarity: model preference is highly correlated with camera geometry, and multi-model fusion brings the largest gains on difficult samples where individual experts are unreliable. Motivated by these observations, we propose \textbf{\ours}, a vision-language agent for adaptive monocular depth estimation. DepthAgent treats existing depth models as frozen tools and learns to analyze scene and camera cues, invoke suitable experts through multi-turn tool utilization, and select or fuse their predictions for each input. To optimize such discrete decision-making toward dense geometric quality, we design a multi-reward reinforcement fine-tuning scheme that jointly encourages valid tool execution, camera/scene analysis, expert-selection quality, and inference efficiency. Extensive experiments across perspective, fisheye, and panoramic benchmarks show that \ours consistently outperforms individual experts, fixed model fusion, and different selection strategies, with strong improvements on challenging samples, highlighting the critical role of expert selection and fusion. The code and model will be released upon publication.

📄 PDF Abstract BibTeX arXiv:2605.23281

Code (0)

등록된 구현이 없습니다.

Tasks

Monocular Depth Estimation

Similar Papers 제목 키워드 기반

UniDAC: Universal Metric Depth Estimation for Any Camera

2026-03-28 · Girish Chandar Ganesan, Yuliang Guo, Liu Ren, Xiaoming Liu arxiv

Monocular metric depth estimation (MMDE) is a core challenge in computer vision, playing a pivotal role in real-world applications that demand accurate spatial understanding. Although prior works have shown promising zer…

Depth Estimation

Adversarial Attacks on Monocular Depth Estimation

2020-03-23 · Ziqi Zhang, Xinge Zhu, Yingwei Li, Xiangqun Chen 외

Recent advances of deep learning have brought exceptional performance on many computer vision tasks such as semantic segmentation and depth estimation. However, the vulnerability of deep neural networks towards adversari…

Autonomous DrivingDepth EstimationMonocular Depth EstimationRobot Navigation+2

SM4Depth: Seamless Monocular Metric Depth Estimation across Multiple Cameras and Scenes by One Model

2024-03-13 · Yihao Liu, Feng Xue, Anlong Ming, Mingshuai Zhao 외

In the last year, universal monocular metric depth estimation (universal MMDE) has gained considerable attention, serving as the foundation model for various multimedia tasks, such as video and image editing. Nonetheless…

Depth EstimationGPU

Uncertainty Guided Depth Fusion for Spike Camera

2022-08-26 · Jianing Li, Jiaming Liu, Xiaobao Wei, Jiyuan Zhang 외

Depth estimation is essential for various important real-world applications such as autonomous driving. However, it suffers from severe performance degradation in high-velocity scenario since traditional cameras can only…

Autonomous DrivingDepth EstimationStereo Depth Estimation

Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure Priors

2025-03-26 · CVPR 2025 1 · Weilong Yan, Ming Li, Haipeng Li, Shuwei Shao 외

Self-supervised depth estimation from monocular cameras in diverse outdoor conditions, such as daytime, rain, and nighttime, is challenging due to the difficulty of learning universal representations and the severe lack …

Depth EstimationWorld Knowledge