paper-with-me

홈 › Papers

Depth Anything 3: Recovering the Visual Space from Any Views

2025-11-13 · Haotong Lin, Sili Chen, Junhao Liew, Donny Y. Chen, Zhenyu Li, Guang Shi, Jiashi Feng, Bingyi Kang arxiv

We present Depth Anything 3 (DA3), a model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known camera poses. In pursuit of minimal modeling, DA3 yields two key insights: a single plain transformer (e.g., vanilla DINO encoder) is sufficient as a backbone without architectural specialization, and a singular depth-ray prediction target obviates the need for complex multi-task learning. Through our teacher-student training paradigm, the model achieves a level of detail and generalization on par with Depth Anything 2 (DA2). We establish a new visual geometry benchmark covering camera pose estimation, any-view geometry and visual rendering. On this benchmark, DA3 sets a new state-of-the-art across all tasks, surpassing prior SOTA VGGT by an average of 44.3% in camera pose accuracy and 25.1% in geometric accuracy. Moreover, it outperforms DA2 in monocular depth estimation. All models are trained exclusively on public academic datasets.

📄 PDF Abstract BibTeX arXiv:2511.10647

Code (0)

등록된 구현이 없습니다.

Tasks

Monocular Depth EstimationCamera Pose EstimationMulti-Task Learning

Similar Papers 제목 키워드 기반

URDF-Anything+: End-to-End Generation for Simulation-Ready Articulated Assets

2026-03-14 · Zhuangzhe Wu, Yue Xin, Chengkai Hou, Minghao Chen 외 arxiv

Articulated objects are fundamental for robotics, simulation of physics, and interactive virtual environments. However, recovering them from visual observations is inherently challenging, as images provide only partial a…

Composition Vision-Language Understanding via Segment and Depth Anything Model

2024-06-07 · Mingxiao Huo, Pengliang Ji, Haotian Lin, Junchen Liu 외

We introduce a pioneering unified library that leverages depth anything, segment anything models to augment neural comprehension in language-vision model zero-shot understanding. This library synergizes the capabilities …

Question AnsweringVisual Question Answering (VQA)

VGLD: Visually-Guided Linguistic Disambiguation for Monocular Depth Scale Recovery

2025-05-05 · Bojin Wu, Jing Chen

We propose a robust method for monocular depth scale recovery. Monocular depth estimation can be divided into two main directions: (1) relative depth estimation, which provides normalized or inverse depth without scale i…

Depth EstimationMonocular Depth Estimation

MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources

2026-01-29 · Baorui Ma, Jiahui Yang, Donglin Di, Xuancheng Zhang 외 arxiv

Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging due to heterogeneous sensor noise, camera-dependent biases, and metric ambiguity…

Monocular Depth EstimationSpatial Reasoning3D ReconstructionDepth Completion

Inter-View Depth Consistency Testing in Depth Difference Subspace

2023-01-27 · Pravin Kumar Rana, Markus Flierl

Multiview depth imagery will play a critical role in free-viewpoint television. This technology requires high quality virtual view synthesis to enable viewers to move freely in a dynamic real world scene. Depth imagery a…

Stereo Matching