paper-with-me

Papers

JEPADepth: Masked Predictive Representation Learning for Self-Supervised Monocular Depth Estimation

2026-07-29 · Ionuţ Grigore, Călin-Adrian Popa arxiv

Self-supervised monocular depth estimation typically relies on photometric reconstruction losses that couple depth, pose, and appearance assumptions. In this paper, we propose JEPADepth, a self-supervised monocular depth framework that incorporates a complementary training objective inspired by Image Joint-Embedding Predictive Architectures (I-JEPA) for self-supervised depth learning. Our method augments a standard photometric pipeline with a masked prediction loss computed in the representation space of a pretrained DINOv3 Vision Transformer encoder. A predictor infers target-region embeddings from visible context-region embeddings under structured masking, and is discarded along with the target encoder at inference time, adding no deployment cost. On KITTI, adding the JEPA objective consistently improves performance over the same DINOv3-based photometric baseline, without changing the inference-time architecture. Compared to prior monocular self-supervised methods, JEPADepth is competitive with state-of-the-art transformer-based approaches and outperforms strong CNN-based baselines on the standard benchmark. In zero-shot transfer (trained on KITTI and evaluated without fine-tuning), JEPADepth achieves the best or near-best performance among the compared methods on both Make3D and Cityscapes across multiple metrics.

📄 PDF Abstract BibTeX arXiv:2607.26600

Code (0)

등록된 구현이 없습니다.

Tasks

Monocular Depth EstimationRepresentation Learning

Similar Papers 제목 키워드 기반

Self-Supervised Learning for Visual Relationship Detection through Masked Bounding Box Reconstruction

2023-11-08 · Zacharias Anastasakis, Dimitrios Mallis, Markos Diomataris, George Alexandridis 외

We present a novel self-supervised approach for representation learning, particularly for the task of Visual Relationship Detection (VRD). Motivated by the effectiveness of Masked Image Modeling (MIM), we propose Masked …

Predicate DetectionRelationship DetectionRepresentation LearningSelf-Supervised Learning+1

Environment Predictive Coding for Visual Navigation

2021-09-29 · ICLR 2022 4 · Santhosh Kumar Ramakrishnan, Tushar Nagarajan, Ziad Al-Halah, Kristen Grauman

We introduce environment predictive coding, a self-supervised approach to learn environment-level representations for embodied agents. In contrast to prior work on self-supervised learning for individual images, we aim t…

Representation LearningSelf-Supervised LearningVisual Navigation

Momentum-Guided Semantic Forecasting (MoFore) for Self-Supervised Video Representation Learning

2026-06-08 · Qinwu Xu arxiv

Self-supervised video representation learning has recently advanced through contrastive learning, masked reconstruction, and predictive representation learning. Reconstruction-based approaches such as MAE and VideoMAE le…

Representation LearningContrastive Learning

Learning and Leveraging World Models in Visual Representation Learning

2024-03-01 · Quentin Garrido, Mahmoud Assran, Nicolas Ballas, Adrien Bardes 외

Joint-Embedding Predictive Architecture (JEPA) has emerged as a promising self-supervised approach that learns by leveraging a world model. While previously limited to predicting missing parts of an input, we explore how…

Representation Learning

HiMAE: Hierarchical Masked Autoencoders Discover Resolution-Specific Structure in Wearable Time Series

2025-10-28 · Simon A. Lee, Cyrus Tanade, Hao Zhou, Juhyeon Lee 외 arxiv

Wearable sensors provide abundant physiological time series, yet the principles governing their predictive utility remain unclear. We hypothesize that temporal resolution is a fundamental axis of representation learning,…

Representation Learning