paper-with-me

Papers

VJEPA: Variational Joint Embedding Predictive Architectures as Probabilistic World Models

2026-01-20 · Yongchao Huang arxiv

Joint Embedding Predictive Architectures (JEPA) offer a scalable paradigm for self-supervised learning by predicting latent representations rather than reconstructing high-entropy observations. However, existing formulations rely on \textit{deterministic} regression objectives, which mask probabilistic semantics and limit its applicability in stochastic control. In this work, we introduce \emph{Variational JEPA (VJEPA)}, a \textit{probabilistic} generalization that learns a predictive distribution over future latent states via a variational objective. We show that VJEPA unifies representation learning with Predictive State Representations (PSRs) and Bayesian filtering, establishing that sequential modeling does not require autoregressive observation likelihoods. Theoretically, we prove that VJEPA representations can serve as sufficient information states for optimal control without pixel reconstruction, while providing formal guarantees for collapse avoidance. We further propose \emph{Bayesian JEPA (BJEPA)}, an extension that factorizes the predictive belief into a learned dynamics expert and a modular prior expert, enabling zero-shot task transfer and constraint (e.g. goal, physics) satisfaction via a Product of Experts. Empirically, through a noisy environment experiment, we demonstrate that VJEPA and BJEPA successfully filter out high-variance nuisance distractors that cause representation collapse in generative baselines. By enabling principled uncertainty estimation (e.g. constructing credible intervals via sampling) while remaining likelihood-free regarding observations, VJEPA provides a foundational framework for scalable, robust, uncertainty-aware planning in high-dimensional, noisy environments.

📄 PDF Abstract BibTeX arXiv:2601.14354

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningRepresentation Learning

Similar Papers 제목 키워드 기반

From Video to EEG: Adapting Joint Embedding Predictive Architecture to Uncover Saptiotemporal Dynamics in Brain Signal Analysis

2025-07-04 · Amirabbas Hojjati, Lu Li, Ibrahim Hameed, Anis Yazidi 외 arxiv

EEG signals capture brain activity with high temporal and low spatial resolution, supporting applications such as neurological diagnosis, cognitive monitoring, and brain-computer interfaces. However, effective analysis i…

Self-Supervised Learning

Video Joint-Embedding Predictive Architectures for Facial Expression Recognition

2026-01-14 · Lennart Eing, Cristina Luna-Jiménez, Silvan Mertes, Elisabeth André arxiv

This paper introduces a novel application of Video Joint-Embedding Predictive Architectures (V-JEPAs) for Facial Expression Recognition (FER). Departing from conventional pre-training methods for video understanding that…

Facial Expression Recognition

Improving the Physics of Video Generation with VJEPA-2 Reward Signal

2025-10-22 · Jianhao Yuan, Xiaofeng Zhang, Felix Friedrich, Nicolas Beltran-Velez 외 arxiv

This is a short technical report describing the winning entry of the PhysicsIQ Challenge, presented at the Perception Test Workshop at ICCV 2025. State-of-the-art video generative models exhibit severely limited physical…

Video Generation

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture

2026-06-22 · Felix Tristram, Stefano Gasperini, Benjamin Killeen, Marcel Walch 외 arxiv

The increasing maturity of embodied AI platforms has driven a growing interest in procedural video representation learning to support intelligent assistance systems for complex, multi-step tasks. Leveraging large-scale l…

Representation LearningAction ClassificationAction Segmentation

FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds

2026-08-02 · Kapil Wanaskar, Gaytri Jena, Aman Chadha, Vinija Jain 외 arxiv

World models have attracted significant attention for their ability to capture and predict the structure and dynamics of the physical world. In this emerging landscape, Joint Embedding Predictive Architectures (JEPA) off…