paper-with-me

홈 › Papers

PureCLIP-Depth: Prompt-Free and Decoder-Free Monocular Depth Estimation within CLIP Embedding Space

2026-03-17 · Ryutaro Miya, Kazuyoshi Fushinobu, Tatsuya Kawaguchi arxiv

We propose PureCLIP-Depth, a completely prompt-free, decoder-free Monocular Depth Estimation (MDE) model that operates entirely within the Contrastive Language-Image Pre-training (CLIP) embedding space. Unlike recent models that rely heavily on geometric features, we explore a novel approach to MDE driven by conceptual information, performing computations directly within the conceptual CLIP space. The core of our method lies in learning a direct mapping from the RGB domain to the depth domain strictly inside this embedding space. Our approach achieves state-of-the-art performance among CLIP embedding-based models on both indoor and outdoor datasets. The code used in this research is available at: https://github.com/ryutaroLF/PureCLIP-Depth

📄 PDF Abstract BibTeX arXiv:2603.16238

Code (0)

등록된 구현이 없습니다.

Tasks

Monocular Depth Estimation

Similar Papers 제목 키워드 기반

MirrorSAM2: Segment Mirror in Videos with Depth Perception

2025-09-21 · Mingchen Xu, Yukun Lai, Ze Ji, Jing Wu arxiv

This paper presents MirrorSAM2, the first framework that adapts Segment Anything Model 2 (SAM2) to the task of RGB-D video mirror segmentation. MirrorSAM2 addresses key challenges in mirror detection, such as reflection …

SPACE-CLIP: Spatial Perception via Adaptive CLIP Embeddings for Monocular Depth Estimation

2026-01-25 · Taewan Cho, Taeryang Kim, Andrew Jaeyong Choi arxiv

Robotic and autonomous systems need dense spatial cues, but many monocular depth models are heavy, task-specific, or hard to attach to an existing multimodal stack. CLIP offers strong semantic representations, yet most C…

Monocular Depth Estimation

SelfPromer: Self-Prompt Dehazing Transformers with Depth-Consistency

2023-03-13 · Cong Wang, Jinshan Pan, WanYu Lin, Jiangxin Dong 외

This work presents an effective depth-consistency self-prompt Transformer for image dehazing. It is motivated by an observation that the estimated depths of an image with haze residuals and its clear counterpart vary. En…

Image DehazingImage Generation

ByDeWay: Boost Your multimodal LLM with DEpth prompting in a Training-Free Way

2025-07-11 · Rajarshi Roy, Devleena Das, Ankesh Banerjee, Arjya Bhattacharjee 외

We introduce ByDeWay, a training-free framework designed to enhance the performance of Multimodal Large Language Models (MLLMs). ByDeWay uses a novel prompting strategy called Layered-Depth-Based Prompting (LDP), which i…

Depth EstimationHallucinationLanguage ModelingLanguage Modelling+2

FreeStyle: Free Lunch for Text-guided Style Transfer using Diffusion Models

2024-01-28 · Feihong He, Gang Li, Fuhui Sun, Mengyuan Zhang 외

The rapid development of generative diffusion models has significantly advanced the field of style transfer. However, most current style transfer methods based on diffusion models typically involve a slow iterative optim…

DecoderStyle Transfer