paper-with-me

홈 › Papers

DMT-JEPA: Discriminative Masked Targets for Joint-Embedding Predictive Architecture

2024-05-28 · Shentong Mo, Sukmin Yun

The joint-embedding predictive architecture (JEPA) recently has shown impressive results in extracting visual representations from unlabeled imagery under a masking strategy. However, we reveal its disadvantages, notably its insufficient understanding of local semantics. This deficiency originates from masked modeling in the embedding space, resulting in a reduction of discriminative power and can even lead to the neglect of critical local semantics. To bridge this gap, we introduce DMT-JEPA, a novel masked modeling objective rooted in JEPA, specifically designed to generate discriminative latent targets from neighboring information. Our key idea is simple: we consider a set of semantically similar neighboring patches as a target of a masked patch. To be specific, the proposed DMT-JEPA (a) computes feature similarities between each masked patch and its corresponding neighboring patches to select patches having semantically meaningful relations, and (b) employs lightweight cross-attention heads to aggregate features of neighboring patches as the masked targets. Consequently, DMT-JEPA demonstrates strong discriminative power, offering benefits across a diverse spectrum of downstream tasks. Through extensive experiments, we demonstrate our effectiveness across various visual benchmarks, including ImageNet-1K image classification, ADE20K semantic segmentation, and COCO object detection tasks. Code is available at: \url{https://github.com/DMTJEPA/DMTJEPA}.

📄 PDF Abstract BibTeX arXiv:2405.17995

Code (1)

dmtjepa/dmtjepa 공식 구현 pytorch

Tasks

image-classificationImage Classificationobject-detectionObject DetectionSemantic Segmentation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture

2025-11-21 · Xiangteng He, Shunsuke Sakai, Shivam Chandhok, Sara Beery 외 arxiv

Recent advances in self-supervised visual representation learning have demonstrated the effectiveness of predictive latent-space objectives for learning transferable features. In particular, Image-based Joint-Embedding P…

Self-Supervised LearningRepresentation LearningImage Classification

Enhancing JEPAs with Spatial Conditioning: Robust and Efficient Representation Learning

2024-10-14 · Etai Littwin, Vimal Thilak, Anand Gopalakrishnan

Image-based Joint-Embedding Predictive Architecture (IJEPA) offers an attractive alternative to Masked Autoencoder (MAE) for representation learning using the Masked Image Modeling framework. IJEPA drives representations…

image-classificationImage ClassificationRepresentation Learning

US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound

2026-02-22 · Ashwath Radhachandran, Vedrana Ivezić, Shreeram Athreya, Ronit Anilkumar 외 arxiv

Ultrasound (US) imaging poses unique challenges for representation learning due to its inherently noisy acquisition process. The low signal-to-noise ratio and stochastic speckle patterns hinder standard self-supervised l…

Self-Supervised LearningRepresentation Learning

HP-JEPA: Hierarchical Partitioning for Multi-Resolution Graph Joint-Embedding Predictive Learning

2026-08-01 · Ruichen Xu, Jingxiang Qu, Wenhan Gao, Jiaxing Zhang 외 arxiv

Graph self-supervised learning aims to learn transferable representations from large-scale unlabeled graph data. Joint-embedding predictive architectures (JEPAs) avoid explicit negative-pair construction and raw-input re…

Graph Representation LearningSelf-Supervised LearningGraph ClassificationGraph Regression

Video Joint-Embedding Predictive Architectures for Facial Expression Recognition

2026-01-14 · Lennart Eing, Cristina Luna-Jiménez, Silvan Mertes, Elisabeth André arxiv

This paper introduces a novel application of Video Joint-Embedding Predictive Architectures (V-JEPAs) for Facial Expression Recognition (FER). Departing from conventional pre-training methods for video understanding that…

Facial Expression Recognition