paper-with-me

홈 › Papers

Why and How Auxiliary Tasks Improve JEPA Representations

2025-09-12 · Jiacan Yu, Siyi Chen, Mingrui Liu, Nono Horiuchi, Vladimir Braverman, Zicheng Xu, Dan Haramati, Randall Balestriero arxiv

Joint-Embedding Predictive Architecture (JEPA) is increasingly used for visual representation learning and as a component in model-based RL, but its behavior remains poorly understood. We provide a theoretical characterization of a simple, practical JEPA variant that has an auxiliary regression head trained jointly with latent dynamics. We prove a No Unhealthy Representation Collapse theorem: in deterministic MDPs, if training drives both the latent-transition consistency loss and the auxiliary regression loss to zero, then any pair of non-equivalent observations, i.e., those that do not have the same transition dynamics or auxiliary value, must map to distinct latent representations. Thus, the auxiliary task anchors which distinctions the representation must preserve. Controlled ablations in a counting environment corroborate the theory and show that training the JEPA model jointly with the auxiliary head generates a richer representation than training them separately. Our work indicates a path to improve JEPA encoders: training them with an auxiliary function that, together with the transition dynamics, encodes the right equivalence relations.

📄 PDF Abstract BibTeX arXiv:2509.12249

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Factorized Latent Dynamics for Video JEPA: An Empirical Study of Auxiliary Objectives

2026-05-16 · Santosh Premi arxiv

Joint-Embedding Predictive Architectures (JEPA) are a promising framework for self-supervised video representation learning, yet the behavior of auxiliary objectives in small-scale Video-JEPA training is not well charact…

Representation Learning

Self-Supervised Ultrasound-Video Segmentation with Feature Prediction and 3D Localised Loss

2025-07-24 · Edward Ellis, Robert Mendel, Andrew Bulpitt, Nasim Parsa 외 arxiv

Acquiring and annotating large datasets in ultrasound imaging is challenging due to low contrast, high noise, and susceptibility to artefacts. This process requires significant time and clinical expertise. Self-supervise…

Self-Supervised LearningVideo Segmentation

TD-JEPA: Latent-predictive Representations for Zero-Shot Reinforcement Learning

2025-10-01 · Marco Bagatella, Matteo Pirotta, Ahmed Touati, Alessandro Lazaric 외 arxiv

Latent prediction--where agents learn by predicting their own latents--has emerged as a powerful paradigm for training general representations in machine learning. In reinforcement learning (RL), this approach has been e…

Reinforcement Learning

Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures

2026-01-30 · Georgios Ioannides, Adrian Kieback, Judah Goldfeder, Linsey Pang 외 arxiv

Joint Embedding Predictive Architectures (JEPA) offer a promising approach to self-supervised speech representation learning, but suffer from representation collapse without explicit grounding. We propose GMM-Anchored JE…

Representation LearningEmotion RecognitionSlot Filling

Representation Without Reward: A JEPA Audit for LLM Fine-Tuning

2026-05-14 · Biswa Sengupta arxiv

Joint-embedding predictive architectures (JEPAs) propose that a model should learn more useful abstractions when trained to predict latent representations rather than observed outputs. For autoregressive language-model f…