Papers Monocular Depth Estimation
“Monocular Depth Estimation” 태그가 달린 논문 1,073편 · 필터 해제
Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation
Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computational photography, and robotics, among others. Despite the field's matur…
Surface Normals EstimationMonocular Depth EstimationImage GenerationWeather-Conditioned Depth Anything
Monocular depth estimation foundation models, such as the Depth Anything series, have achieved remarkable performance across diverse domains. However, they still suffer from critical failures under adverse weather condit…
Monocular Depth EstimationEfficient and High-Quality Depth Estimation via Pixel-Space Diffusion with Linear Attention
This work presents $\textbf{Lapis}$, a $\textbf{l}$inear-$\textbf{a}$ttention-based $\textbf{pi}$xel-$\textbf{s}$pace generative framework that achieves efficient and high-fidelity depth estimation with one-step diffusio…
Monocular Depth EstimationFrom Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation
Vision foundation models are capable of generalizing across 3-dimensional (3D) scenes with high-fidelity estimates; their empirical success can be attributed to training on large-scale datasets of perspective images. How…
Monocular Depth EstimationBinarized High-Efficiency RAW Video Restoration and Beyond
RAW video restoration is fundamental to high-quality low-level perception and serves as the basis for a wide range of downstream vision applications. While binary neural networks (BNNs) enable efficient lightweight deplo…
Monocular Depth EstimationVideo RestorationImage EnhancementObject DetectionPXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation
Recent monocular depth estimators achieve strong zero-shot generalization, yet often struggle to preserve fine-grained structures and object boundaries. We attribute this limitation to the prevalent combination of large-…
Monocular Depth EstimationZero-shot GeneralizationXiDepth: a Lightweight and Efficient Network for Self-supervised Monocular Depth Estimation
Self-supervised monocular depth estimation has emerged as an appealing solution to design lightweight and effective models for deployment on computationally constrained devices due to its reduced reliance on expensive de…
Monocular Depth EstimationBreaking the Horizontal Prior: From Long-Tailed Orientation Bias to Roll-Robust Monocular Depth Estimation
Despite recent advances in Monocular Depth Estimation, state-of-the-art depth foundation models remain vulnerable to robustness issues. Particularly, even slight camera rolls can result in substantial degradation in dept…
Monocular Depth EstimationData AugmentationSpatial ReasoningBeyond Visual Ambiguity: Guiding Robust Monocular Depth Estimation in Challenging Scenarios via Detailed Long Captions
Monocular depth estimation (MDE) faces challenges with non-Lambertian surfaces and adverse weather conditions due to the visual ambiguities inherent in single-image limited information. Existing works address them in iso…
Monocular Depth EstimationImage InpaintingJEPADepth: Masked Predictive Representation Learning for Self-Supervised Monocular Depth Estimation
Self-supervised monocular depth estimation typically relies on photometric reconstruction losses that couple depth, pose, and appearance assumptions. In this paper, we propose JEPADepth, a self-supervised monocular depth…
Monocular Depth EstimationRepresentation LearningDAPM: UAV Monocular Depth Estimation from Any Height, Pitch, Roll and FOV
Monocular depth estimation is a fundamental prerequisite for 3D reconstruction and autonomous navigation in Unmanned Aerial Vehicles (UAVs). In practical deployments, UAVs operate under highly dynamic camera poses charac…
Monocular Depth Estimation3D ReconstructionPose EstimationWAT3R: Feedforward Underwater 3D Reconstruction
Reliable feedforward underwater 3D reconstruction remains challenging due to severe light attenuation and backscattering, which degrade visual quality and disrupt feature consistency across views, leading to inaccurate m…
Monocular Depth EstimationCamera Pose Estimation3D ReconstructionARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning
Diffusion models have recently become the dominant paradigm for monocular depth estimation (MDE). However, they implicitly assume that depth can be recovered as a globally smooth field through iterative denoising, which …
Monocular Depth EstimationZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device
Monocular depth estimation has seen remarkable progress through foundation models achieving robust zero-shot generalization, yet their computational demands place them far beyond the reach of embedded and mobile platform…
Monocular Depth EstimationZero-shot GeneralizationKnowledge DistillationTime-to-Collision Based Dynamic Obstacle Avoidance Using Pretrained Vision Models for Robots in Unstructured Environments
Dynamic obstacle avoidance in unstructured outdoor environments remains a critical challenge for autonomous mobile robots, particularly when large-scale robot-specific training data and simulation-based policies are impr…
Monocular Depth EstimationMonocular Vision Based Control Framework for Grasping
Grasping in unstructured environments requires handling objects with widely different mechanical properties, from soft and deformable items to rigid everyday objects. Most existing approaches address these categories sep…
Monocular Depth EstimationImage SegmentationObject DetectionPoint TrackingA Task-Driven Evaluation of UAV Detection and Tracking under Synthetic Fog
Fog severely degrades the visibility of small unmanned aerial vehicles (UAVs) in skydominant, long-range imagery, reducing the reliability of downstream detection and tracking. This paper presents a task-driven evaluatio…
Monocular Depth EstimationImage RestorationObject DetectionTRISTAR: Triple-Signal Stair Recognition and Vision-Only Indoor Navigation for Search-and-Rescue Micro-UAVs
Indoor search-and-rescue (SAR) operations often require rapid situational awareness where GNSS signals are unavailable and human access is difficult or hazardous. While most autonomous aerial systems rely on LiDAR, stere…
Monocular Depth EstimationScene UnderstandingVisionAId: An Offline-First Multimodal Android Assistant for People with Visual Impairment, Featuring Personalized Object Retrieval
Over 285 million people worldwide live with a visual impairment, for whom everyday tasks such as avoiding obstacles, locating personal belongings, recognizing familiar faces, or handling cash remain persistent obstacles …
Monocular Depth EstimationInstance SegmentationSpeech SynthesisFace DetectionActive Spatial Guidance: Eliminating Injected Positional Mechanisms in Vision Transformers
Vision Transformers (ViTs) commonly rely on injected positional mechanisms to address self-attention's permutation invariance. Motivated by the spatial regularities of natural images, we ask whether spatial organization …
Monocular Depth EstimationSemantic Segmentation