Vision-Language Embodiment for Monocular Depth Estimation
Depth estimation is a core problem in robotic perception and vision tasks, but 3D reconstruction from a single image presents inherent uncertainties. Current depth estimation models primarily rely on inter-image relationships for supervised training, often overlooking the intrinsic information provided by the camera itself. We propose a method that embodies the camera model and its physical characteristics into a deep learning model, computing embodied scene depth through real-time interactions with road environments. The model can calculate embodied scene depth in real-time based on immediate environmental changes using only the intrinsic properties of the camera, without any additional equipment. By combining embodied scene depth with RGB image features, the model gains a comprehensive perspective on both geometric and visual details. Additionally, we incorporate text descriptions containing environmental content and depth information as priors for scene understanding, enriching the model's perception of objects. This integration of image and language -- two inherently ambiguous modalities -- leverages their complementary strengths for monocular depth estimation. The real-time nature of the embodied language and depth prior model ensures that the model can continuously adjust its perception and behavior in dynamic environments. Experimental results show that the embodied depth estimation method enhances model performance across different scenes.
Code (0)
등록된 구현이 없습니다.
Tasks
3D ReconstructionDepth EstimationMonocular Depth EstimationScene UnderstandingSimilar Papers 제목 키워드 기반
Embodiment: Self-Supervised Depth Estimation Based on Camera Models
Depth estimation is a critical topic for robotics and vision-related tasks. In monocular depth estimation, in comparison with supervised learning that requires expensive ground truth labeling, self-supervised methods pos…
3D ReconstructionDepth EstimationMonocular Depth EstimationSelf-Supervised LearningCeRLP: A Cross-embodiment Robot Local Planning Framework for Visual Navigation
Visual navigation for cross-embodiment robots is challenging due to variations in robot and camera configurations, which can lead to the failure of navigation tasks. Previous approaches typically rely on collecting massi…
Vision-Language NavigationMonocular Depth EstimationVisual NavigationLarge Language Models Can Understanding Depth from Monocular Images
Monocular depth estimation is a critical function in computer vision applications. This paper shows that large language models (LLMs) can effectively interpret depth with minimal supervision, using efficient resource uti…
Depth EstimationMonocular Depth Estimation3D Visual Illusion Depth Estimation
3D visual illusion is a perceptual phenomenon where a two-dimensional plane is manipulated to simulate three-dimensional spatial relationships, making a flat artwork or object look three-dimensional in the human visual s…
Common Sense ReasoningDepth EstimationLanguage ModelingLanguage ModellingFIS-Nets: Full-image Supervised Networks for Monocular Depth Estimation
This paper addresses the importance of full-image supervision for monocular depth estimation. We propose a semi-supervised architecture, which combines both unsupervised framework of using image consistency and supervise…
Depth CompletionDepth EstimationMonocular Depth Estimation