paper-with-me

Papers Monocular Depth Estimation

“Monocular Depth Estimation” 태그가 달린 논문 1,073편 · 필터 해제

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

2026-09-08 · Igor Pavlovic, Thiemo Wandel, Anton Obukhov, Luca Bartolomei 외 hf

Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computational photography, and robotics, among others. Despite the field's matur…

Surface Normals EstimationMonocular Depth EstimationImage Generation

Weather-Conditioned Depth Anything

2026-09-04 · Zhaoming Xu, Chan-Wei Hu, Kuan-Ru Huang, Zihao Zhu 외 arxiv

Monocular depth estimation foundation models, such as the Depth Anything series, have achieved remarkable performance across diverse domains. However, they still suffer from critical failures under adverse weather condit…

Monocular Depth Estimation

Efficient and High-Quality Depth Estimation via Pixel-Space Diffusion with Linear Attention

2026-08-31 · Bingde Liu, Wu Ran, Jinglei Zhang, Huanhuan Yuan 외 arxiv

This work presents $\textbf{Lapis}$, a $\textbf{l}$inear-$\textbf{a}$ttention-based $\textbf{pi}$xel-$\textbf{s}$pace generative framework that achieves efficient and high-fidelity depth estimation with one-step diffusio…

Monocular Depth Estimation

From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation

2026-08-28 · Rit Gangopadhyay, Alex Wong arxiv

Vision foundation models are capable of generalizing across 3-dimensional (3D) scenes with high-fidelity estimates; their empirical success can be attributed to training on large-scale datasets of perspective images. How…

Monocular Depth Estimation

Binarized High-Efficiency RAW Video Restoration and Beyond

2026-08-17 · Tianyu Zhu, Ying Fu, Hesong Li, Gengchen Zhang 외 arxiv

RAW video restoration is fundamental to high-quality low-level perception and serves as the basis for a wide range of downstream vision applications. While binary neural networks (BNNs) enable efficient lightweight deplo…

Monocular Depth EstimationVideo RestorationImage EnhancementObject Detection

PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation

2026-08-17 · Zhiyuan Yuan, Guanying Chen, Lingteng Qiu, Ruimao Zhang 외 arxiv

Recent monocular depth estimators achieve strong zero-shot generalization, yet often struggle to preserve fine-grained structures and object boundaries. We attribute this limitation to the prevalent combination of large-…

Monocular Depth EstimationZero-shot Generalization

XiDepth: a Lightweight and Efficient Network for Self-supervised Monocular Depth Estimation

2026-08-04 · Elena Izzo, Riccardo Toniolo, Lamberto Ballan arxiv

Self-supervised monocular depth estimation has emerged as an appealing solution to design lightweight and effective models for deployment on computationally constrained devices due to its reduced reliance on expensive de…

Monocular Depth Estimation

Breaking the Horizontal Prior: From Long-Tailed Orientation Bias to Roll-Robust Monocular Depth Estimation

2026-08-01 · Kaihua Tang, Ziqing Xia, Xiaoxu Zheng, Xiaoxue Zhang 외 arxiv

Despite recent advances in Monocular Depth Estimation, state-of-the-art depth foundation models remain vulnerable to robustness issues. Particularly, even slight camera rolls can result in substantial degradation in dept…

Monocular Depth EstimationData AugmentationSpatial Reasoning

Beyond Visual Ambiguity: Guiding Robust Monocular Depth Estimation in Challenging Scenarios via Detailed Long Captions

2026-07-30 · Junrui Zhang, Jiaqi Li, Yiran Wang, Liao Shen 외 arxiv

Monocular depth estimation (MDE) faces challenges with non-Lambertian surfaces and adverse weather conditions due to the visual ambiguities inherent in single-image limited information. Existing works address them in iso…

Monocular Depth EstimationImage Inpainting

JEPADepth: Masked Predictive Representation Learning for Self-Supervised Monocular Depth Estimation

2026-07-29 · Ionuţ Grigore, Călin-Adrian Popa arxiv

Self-supervised monocular depth estimation typically relies on photometric reconstruction losses that couple depth, pose, and appearance assumptions. In this paper, we propose JEPADepth, a self-supervised monocular depth…

Monocular Depth EstimationRepresentation Learning

DAPM: UAV Monocular Depth Estimation from Any Height, Pitch, Roll and FOV

2026-07-23 · Tong Ling, Wenhui Diao, Yingchao Feng, Hanbo Bi 외 arxiv

Monocular depth estimation is a fundamental prerequisite for 3D reconstruction and autonomous navigation in Unmanned Aerial Vehicles (UAVs). In practical deployments, UAVs operate under highly dynamic camera poses charac…

Monocular Depth Estimation3D ReconstructionPose Estimation

WAT3R: Feedforward Underwater 3D Reconstruction

2026-07-23 · Jiayi Xu, Jiahao Lu, Ziqiang Zheng, Yihao Tan 외 arxiv

Reliable feedforward underwater 3D reconstruction remains challenging due to severe light attenuation and backscattering, which degrade visual quality and disrupt feature consistency across views, leading to inaccurate m…

Monocular Depth EstimationCamera Pose Estimation3D Reconstruction

ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning

2026-07-14 · Zijie Wang, Wei Zhang, Weiming Zhang, Xiao Tan 외 arxiv

Diffusion models have recently become the dominant paradigm for monocular depth estimation (MDE). However, they implicitly assume that depth can be recovered as a globally smooth field through iterative denoising, which …

Monocular Depth Estimation

ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device

2026-07-09 · Fabio Tosi, Luca Bartolomei, Matteo Poggi, Stefano Mattoccia arxiv

Monocular depth estimation has seen remarkable progress through foundation models achieving robust zero-shot generalization, yet their computational demands place them far beyond the reach of embedded and mobile platform…

Monocular Depth EstimationZero-shot GeneralizationKnowledge Distillation

Time-to-Collision Based Dynamic Obstacle Avoidance Using Pretrained Vision Models for Robots in Unstructured Environments

2026-07-08 · Erik Jagnandan, Mulugeta Haile, Gregory Barber, Pratik Chaudhari arxiv

Dynamic obstacle avoidance in unstructured outdoor environments remains a critical challenge for autonomous mobile robots, particularly when large-scale robot-specific training data and simulation-based policies are impr…

Monocular Depth Estimation

Monocular Vision Based Control Framework for Grasping

2026-07-08 · Shail Jadav, Dongheui Lee arxiv

Grasping in unstructured environments requires handling objects with widely different mechanical properties, from soft and deformable items to rigid everyday objects. Most existing approaches address these categories sep…

Monocular Depth EstimationImage SegmentationObject DetectionPoint Tracking

A Task-Driven Evaluation of UAV Detection and Tracking under Synthetic Fog

2026-07-06 · Amir Pouladi, Vesal Ahsani, Haijun Li, Homayoun Najjaran 외 arxiv

Fog severely degrades the visibility of small unmanned aerial vehicles (UAVs) in skydominant, long-range imagery, reducing the reliability of downstream detection and tracking. This paper presents a task-driven evaluatio…

Monocular Depth EstimationImage RestorationObject Detection

TRISTAR: Triple-Signal Stair Recognition and Vision-Only Indoor Navigation for Search-and-Rescue Micro-UAVs

2026-07-04 · Octavian Gîngu, Stelian Spînu arxiv

Indoor search-and-rescue (SAR) operations often require rapid situational awareness where GNSS signals are unavailable and human access is difficult or hazardous. While most autonomous aerial systems rely on LiDAR, stere…

Monocular Depth EstimationScene Understanding

VisionAId: An Offline-First Multimodal Android Assistant for People with Visual Impairment, Featuring Personalized Object Retrieval

2026-07-02 · Cristian-Gabriel Florea, Stelian Spînu arxiv

Over 285 million people worldwide live with a visual impairment, for whom everyday tasks such as avoiding obstacles, locating personal belongings, recognizing familiar faces, or handling cash remain persistent obstacles …

Monocular Depth EstimationInstance SegmentationSpeech SynthesisFace Detection

Active Spatial Guidance: Eliminating Injected Positional Mechanisms in Vision Transformers

2026-07-01 · Cong Liu, Xiaofang Li, Simon X. Yang arxiv

Vision Transformers (ViTs) commonly rely on injected positional mechanisms to address self-attention's permutation invariance. Motivated by the spatial regularities of natural images, we ask whether spatial organization …

Monocular Depth EstimationSemantic Segmentation
1–20 / 1,073 다음 →