paper-with-me

홈 › Papers

CLIP Can Understand Depth

2024-02-05 · Dunam Kim, Seokju Lee

Recent studies on generalizing CLIP for monocular depth estimation reveal that CLIP pre-trained on web-crawled data is inefficient for deriving proper similarities between image patches and depth-related prompts. In this paper, we adapt CLIP for meaningful quality of monocular depth estimation with dense prediction, without fine-tuning its original vision-language alignment. By jointly training a compact deconvolutional decoder with a tiny learnable embedding matrix named mirror, as a static prompt for its text encoder, CLIP is enabled to understand depth. With this approach, our model exhibits impressive performance matching several previous state-of-the-art vision-only models on the NYU Depth v2 and KITTI datasets, outperforming every CLIP-based depth estimation model with a large margin. Experiments on temporal depth consistency and spatial continuity demonstrate that the prior knowledge of CLIP can be effectively refined by our proposed framework. Furthermore, an ablation study on mirror proves that the resulting model estimates depth utilizing knowledge not only from the image encoder but also text encoder despite not being given any prompt written in a human way. This research demonstrates that through minimal adjustments, the prior knowledge of vision-language foundation models, such as CLIP, can be generalized even to domains where learning during pretraining is challenging. We facilitate future works focused on methods to adjust suboptimal prior knowledge of vision-language models using non-human language prompts, achieving performance on par with task-specific state-of-the-art methodologies.

📄 PDF Abstract BibTeX arXiv:2402.03251

Code (0)

등록된 구현이 없습니다.

Tasks

Depth EstimationMonocular Depth Estimation

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Can Language Understand Depth?

2022-07-03 · Renrui Zhang, Ziyao Zeng, Ziyu Guo, Yafeng Li

Besides image classification, Contrastive Language-Image Pre-training (CLIP) has accomplished extraordinary success for a wide range of vision tasks, including object-level and 3D space understanding. However, it's still…

Depth Estimationimage-classificationImage ClassificationMonocular Depth Estimation

CLIP-based Point Cloud Classification via Point Cloud to Image Translation

2024-08-07 · Shuvozit Ghose, Manyi Li, Yiming Qian, Yang Wang

Point cloud understanding is an inherently challenging problem because of the sparse and unordered structure of the point cloud in the 3D space. Recently, Contrastive Vision-Language Pre-training (CLIP) based point cloud…

ClassificationPoint Cloud ClassificationTranslation

CLIP2Point: Transfer CLIP to Point Cloud Classification with Image-Depth Pre-training

2022-10-03 · ICCV 2023 1 · Tianyu Huang, Bowen Dong, Yunhan Yang, Xiaoshui Huang 외

Pre-training across 3D vision and language remains under development because of limited training data. Recent works attempt to transfer vision-language pre-training models to 3D vision. PointCLIP converts point cloud dat…

3D Point Cloud ClassificationContrastive LearningFew-Shot LearningPoint Cloud Classification+4

PureCLIP-Depth: Prompt-Free and Decoder-Free Monocular Depth Estimation within CLIP Embedding Space

2026-03-17 · Ryutaro Miya, Kazuyoshi Fushinobu, Tatsuya Kawaguchi arxiv

We propose PureCLIP-Depth, a completely prompt-free, decoder-free Monocular Depth Estimation (MDE) model that operates entirely within the Contrastive Language-Image Pre-training (CLIP) embedding space. Unlike recent mod…

Monocular Depth Estimation

Do Vision-Language Models Understand Compound Nouns?

2024-03-30 · Sonal Kumar, Sreyan Ghosh, S Sakshi, Utkarsh Tyagi 외

Open-vocabulary vision-language models (VLMs) like CLIP, trained using contrastive loss, have emerged as a promising new paradigm for text-to-image retrieval. However, do VLMs understand compound nouns (CNs) (e.g., lab c…

Image RetrievalLanguage ModelingLanguage ModellingLarge Language Model+1