paper-with-me

Papers

Joint-Embedding Predictive Architecture for Self-Supervised Learning of Mask Classification Architecture

2024-07-15 · Dong-Hee Kim, Sungduk Cho, Hyeonwoo Cho, Chanmin Park, Jinyoung Kim, Won Hwa Kim

In this work, we introduce Mask-JEPA, a self-supervised learning framework tailored for mask classification architectures (MCA), to overcome the traditional constraints associated with training segmentation models. Mask-JEPA combines a Joint Embedding Predictive Architecture with MCA to adeptly capture intricate semantics and precise object boundaries. Our approach addresses two critical challenges in self-supervised learning: 1) extracting comprehensive representations for universal image segmentation from a pixel decoder, and 2) effectively training the transformer decoder. The use of the transformer decoder as a predictor within the JEPA framework allows proficient training in universal image segmentation tasks. Through rigorous evaluations on datasets such as ADE20K, Cityscapes and COCO, Mask-JEPA demonstrates not only competitive results but also exceptional adaptability and robustness across various training scenarios. The architecture-agnostic nature of Mask-JEPA further underscores its versatility, allowing seamless adaptation to various mask classification family.

📄 PDF Abstract BibTeX arXiv:2407.10733

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderImage SegmentationSegmentationSelf-Supervised LearningSemantic Segmentation

Similar Papers 제목 키워드 기반

MC-JEPA: A Joint-Embedding Predictive Architecture for Self-Supervised Learning of Motion and Content Features

2023-07-24 · Adrien Bardes, Jean Ponce, Yann Lecun

Self-supervised learning of visual representations has been focusing on learning content features, which do not capture object motion or location, and focus on identifying and differentiating objects in images and videos…

Optical Flow EstimationSelf-Supervised LearningSemantic Segmentation

AD-L-JEPA: Self-Supervised Spatial World Models with Joint Embedding Predictive Architecture for Autonomous Driving with LiDAR Data

2025-01-09 · Haoran Zhu, Zhenyuan Dong, Kristi Topollai, Anna Choromanska

As opposed to human drivers, current autonomous driving systems still require vast amounts of labeled data to train. Recently, world models have been proposed to simultaneously enhance autonomous driving capabilities by …

3D Object DetectionAutonomous DrivingContrastive LearningTransfer Learning

MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning

2026-06-23 · Revant Teotia, Adrien Bardes, Michael Rabbat, Sumit Chopra 외 arxiv

Self-supervised learning from large-scale video data has emerged as a dominant paradigm for visual representation learning. Since audio and visual streams naturally co-occur in video data, extending this success to joint…

Self-Supervised LearningRepresentation Learning

Point-JEPA: A Joint Embedding Predictive Architecture for Self-Supervised Learning on Point Cloud

2024-04-25 · Ayumu Saito, Prachi Kudeshia, Jiju Poovvancheri

Recent advancements in self-supervised learning in the point cloud domain have demonstrated significant potential. However, these methods often suffer from drawbacks, including lengthy pre-training time, the necessity of…

3D Part Segmentation3D Point Cloud Classification3D Point Cloud Linear ClassificationClassification+2

DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture

2025-11-21 · Xiangteng He, Shunsuke Sakai, Shivam Chandhok, Sara Beery 외 arxiv

Recent advances in self-supervised visual representation learning have demonstrated the effectiveness of predictive latent-space objectives for learning transferable features. In particular, Image-based Joint-Embedding P…

Self-Supervised LearningRepresentation LearningImage Classification