paper-with-me

홈 › Papers

Factorized Latent Dynamics for Video JEPA: An Empirical Study of Auxiliary Objectives

2026-05-16 · Santosh Premi arxiv

Joint-Embedding Predictive Architectures (JEPA) are a promising framework for self-supervised video representation learning, yet the behavior of auxiliary objectives in small-scale Video-JEPA training is not well characterized. We report a small-scale empirical study of 18 auxiliary objective variants for Video-JEPA across two pretraining regimes: single-dataset (UCF-101) and mixed-dataset (UCF-101 + Something-Something V2 + ImageNet-100). We evaluate frozen representations on three complementary benchmarks: Diving-48 (fine-grained motion), SomethingSomething V2 (temporal reasoning), and ImageNet-100 (appearance). Our experiments suggest that many auxiliary objectives exhibit capacity trade-offs: gains on one downstream capability often coincide with degradation on another. We then study FWM-HW-LD (Factorized World-Model with Hard-Region-Weighted Latent Dynamics), a training-time objective that separates the latent representation into appearance and dynamics subspaces and applies hard-region weighting to both JEPA prediction errors and latent dynamics errors. In our mixed-dataset setting, FWM-HW-LD improves ImageNet-100 by +5.92 and SSv2 by +3.21 percentage points relative to the reference baseline, while remaining within 0.30 percentage points on Diving-48. These results indicate that latent factorization is a useful direction for studying auxiliary-objective trade-offs in Video-JEPA.

📄 PDF Abstract BibTeX arXiv:2605.17165

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

JEPA-Anything: Learning Predictive Models across Different Worlds

2026-09-17 · Taoyong Cui, Zhongyao Wang, Xinyue Xu, Weiyang Liu 외 hf

World modeling enables intelligence to anticipate consequences, guide interventions, and learn from interaction. Yet predictive models remain domain-specific: can a common learning principle support world modeling across…

Representation Learning

JEPA-x: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics

2026-08-25 · Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi arxiv

Latent world models plan by predicting how candidate actions advance learned latent dynamics. In self-predictive models, however, the encoder and predictor are optimized jointly and can co-adapt to latent transitions tha…

A Generalization Theory for JEPA-Based World Models

2026-06-25 · Jingyi Cui, Qi Zhang, Hongwei Wen, Yisen Wang arxiv

Joint Embedding Predictive Architectures (JEPAs) have recently emerged as a promising paradigm for world modeling by learning predictive dynamics in a latent space rather than generating future observations at the input …

Graph Learning

Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Control

2026-07-06 · Ruslan Rakhimov, George Bredis, Yuriy Maksyuta, Daniil Gavrilov arxiv

Joint-Embedding Predictive Architectures (JEPAs) underpin a growing family of latent world models for control from raw pixels, but every existing JEPA world model commits at training time to a single inference paradigm: …

VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model

2026-02-10 · Jingwen Sun, Wenyao Zhang, Zekun Qi, Shaojie Ren 외 arxiv

Pretraining Vision-Language-Action (VLA) policies on internet-scale video is appealing, yet current latent-action objectives often learn the wrong thing: they remain anchored to pixel variation rather than action-relevan…