paper-with-me

Papers

Joint Embedding Predictive Architectures Focus on Slow Features

2022-11-20 · Vlad Sobal, Jyothir S V, Siddhartha Jalagam, Nicolas Carion, Kyunghyun Cho, Yann Lecun

Many common methods for learning a world model for pixel-based environments use generative architectures trained with pixel-level reconstruction objectives. Recently proposed Joint Embedding Predictive Architectures (JEPA) offer a reconstruction-free alternative. In this work, we analyze performance of JEPA trained with VICReg and SimCLR objectives in the fully offline setting without access to rewards, and compare the results to the performance of the generative architecture. We test the methods in a simple environment with a moving dot with various background distractors, and probe learned representations for the dot's location. We find that JEPA methods perform on par or better than reconstruction when distractor noise changes every time step, but fail when the noise is fixed. Furthermore, we provide a theoretical explanation for the poor performance of JEPA-based methods with fixed noise, highlighting an important limitation.

📄 PDF Abstract BibTeX arXiv:2211.10831

Code (1)

vladisai/jepa_ssl_neurips_2022 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

Bitcoin Customer Service Number +1-833-534-1729 설명 없음
fail 설명 없음
Test 설명 없음
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Batch Normalization 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Residual Connection 설명 없음
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…

Similar Papers 제목 키워드 기반

Learning Invariant Visual Representations for Planning with Joint-Embedding Predictive World Models

2026-02-20 · Leonardo F. Toso, Davit Shadunts, Yunyang Lu, Nihal Sharma 외 arxiv

World models learned from high-dimensional visual observations allow agents to make decisions and plan directly in latent space, avoiding pixel-level reconstruction. However, recent latent predictive architectures (JEPAs…

Graph-level Representation Learning with Joint-Embedding Predictive Architectures

2023-09-27 · Geri Skenderi, Hang Li, Jiliang Tang, Marco Cristani

Joint-Embedding Predictive Architectures (JEPAs) have recently emerged as a novel and powerful technique for self-supervised representation learning. They aim to learn an energy-based model by predicting the latent repre…

Contrastive LearningData AugmentationGraph ClassificationGraph Regression+1

Investigating Design Choices in Joint-Embedding Predictive Architectures for General Audio Representation Learning

2024-05-14 · Alain Riou, Stefan Lattner, Gaëtan Hadjeres, Geoffroy Peeters

This paper addresses the problem of self-supervised general-purpose audio representation learning. We explore the use of Joint-Embedding Predictive Architectures (JEPA) for this task, which consists of splitting an input…

Audio ClassificationRepresentation Learning

Learning State-Space Models of Dynamic Systems from Arbitrary Data using Joint Embedding Predictive Architectures

2025-08-14 · Jonas Ulmen, Ganesh Sundaram, Daniel Görges arxiv

With the advent of Joint Embedding Predictive Architectures (JEPAs), which appear to be more capable than reconstruction-based methods, this paper introduces a novel technique for creating world models using continuous-t…

Video Joint-Embedding Predictive Architectures for Facial Expression Recognition

2026-01-14 · Lennart Eing, Cristina Luna-Jiménez, Silvan Mertes, Elisabeth André arxiv

This paper introduces a novel application of Video Joint-Embedding Predictive Architectures (V-JEPAs) for Facial Expression Recognition (FER). Departing from conventional pre-training methods for video understanding that…

Facial Expression Recognition