Papers Depth Estimation
“Depth Estimation” 태그가 달린 논문 2,772편 · 필터 해제
DART: Depth-as-Target Pretraining for Surgical Vision Foundation Models
Vision foundation models (VFMs) are valuable in data-scarce domains such as surgery, where a single pretrained backbone can provide rich representations for many downstream tasks. Yet the dominant self-supervised pretrai…
Depth EstimationNeighbor-Aware View Synthesis for Restoring Missing Views in Light-Field Camera Arrays
In light-field (LF) imaging systems, dense spatial sampling from a camera array enables powerful post-capture capabilities such as refocusing and depth estimation. However, real-world LF capture is often affected by hard…
Depth EstimationGeometry-Grounded Unified 3D Perception for Autonomous Driving
Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera streams. However, existing image-based frameworks often rely on backbones pr…
3D Object DetectionAutonomous DrivingDepth EstimationSpatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing work has mainly followed two lines. One lin…
Reinforcement LearningSpatial Reasoning3D ReconstructionDepth EstimationBoosting Generalizable Depth Estimation in Endoscopy by Mixture of Lightweight Experts and Intrinsic Image Alignment
Depth estimation is a significant task for 3D perception in endoscopic surgeries. However, illumination interference and feature diversity in various endoscopic scenes are still challenges for generalizable depth estimat…
parameter-efficient fine-tuningDepth EstimationBoosting Robustness for All-Weather Self-Supervised Depth Estimation in Autonomous Driving
Self-supervised depth estimation is challenging for safe autonomous driving under various adverse weather conditions due to sensor perception degradation. These challenges arise from two main aspects. Firstly, adverse co…
Knowledge DistillationAutonomous DrivingDepth EstimationTransBiolab: A Real-World Multi-View Dataset of Cluttered Transparent Biomedical Objects
Autonomous biomedical laboratories increasingly rely on visual perception to recognize, localize, and manipulate transparent plasticware, yet high-quality real-world datasets for this setting remain limited. The scarcity…
Robot Manipulation6D Pose EstimationDepth EstimationReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video
Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal…
Depth EstimationDPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense Prediction
Multi-Task Learning (MTL) in robotics perception systems supports comprehensive 3D spatial scene understanding by integrating semantic segmentation and depth estimation. While Vision Foundation Models (VFMs) are increasi…
Semantic SegmentationMulti-Task LearningScene UnderstandingDepth EstimationGeometric Distillation from Rectified Stereo: Leveraging Epipolar Cues for Monocular Depth
Monocular depth foundation models have demonstrated remarkable generalization capabilities across diverse environments. However, they continue to struggle with metric depth estimation in diverse environments. This limita…
Depth EstimationLet RGB Be the Language of Vision
This work introduces a unified formulation for vision models, where diverse forms of visual information beyond natural images, such as masks, depth maps, and other structured visual signals, are all represented as RGB im…
Depth EstimationImage GenerationImage EditingOmniX: Any-view and Any-time 4D Reconstruction via Feed-forward Trajectory Fields
Previous feed-forward 4D reconstruction methods either predict per-frame static point clouds, ignoring foreground motion, or estimate point cloud trajectories while being limited to small camera motions. This restricts t…
Camera Pose EstimationTrajectory PredictionDepth EstimationPoint TrackingWat3R: Underwater 3D Geometry Learning without Annotations
Estimating 3D geometry in underwater environments presents unique challenges due to light attenuation, scattering, and the absence of large-scale, high-quality 3D annotations. Pioneering methods rely on massive dense ann…
3D ReconstructionDepth EstimationVision as Unified Multimodal Generation
We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectu…
Camera Pose Estimationmultimodal generationDepth EstimationImage GenerationFrom Foundation to Application: Improving VLA Models in Practice
Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applications continues to impede their practical implementation. To bridge this gap, we present LingBot-VLA 2.0…
Depth EstimationGen4U: Unifying Video Generation and Understanding via Diffusion
Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics. We demonstrate that state-of-the-art video diffusion models overcome this limitation. By systematicall…
Camera Pose EstimationVideo ClassificationDepth EstimationVideo CaptioningVision Pretraining for Dense Spatial Perception
Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable representations from pixel observations. Modern visual foundation models te…
Depth CompletionDepth EstimationOmniDS: Dual-Stream Context Fusion for Omnidirectional Depth from Fisheye Cameras
Omnidirectional depth estimation from multi-fisheye camera rigs is complicated by visibility conflicts: wide baselines cause different cameras to observe different portions, or even different faces, of the same object, s…
Depth EstimationLearning to Suppress SPAD-based LiDAR Flare
Single-Photon Avalanche Diode (SPAD)-based Light Detection and Ranging (LiDAR) is emerging for autonomous vehicles due to its high sensitivity and precise depth sensing capabilities. However, flare caused by excessive ph…
Semantic SegmentationAutonomous VehiclesDepth EstimationPoint CloudsReal-Time Visual Intelligence on Low-Cost UAVs: A Modular Approach for Tracking, Scanning, and Navigation
Autonomous drones are rapidly transforming modern warfare and civil applications alike. This paper presents the development of an integrated intelligent drone system designed to serve as a personal assistant. Leveraging …
Depth Estimation