paper-with-me

홈 › Papers

Subspace-Decomposed JEPAs: Disentangling Progression and Content in Latent World Models

2026-05-29 · Lucas Thil, Jesse Read, Rim Kaddah, Guillaume Doquet arxiv

Joint-Embedding Predictive Architectures (JEPAs) learn compact latent world models by predicting future embeddings, but no single coordinate of the latent is designated to encode task progression. We carve the JEPA latent into two orthogonal subspaces with disjoint roles: a low-dimensional progression subspace shaped by a cosine-margin triplet loss, and a high-dimensional content subspace regularised by the existing SIGReg objective of LeWM. We prove that the two anti-collapse forces act on disjoint coordinates, so they compose additively rather than competing on the same dimensions. Our method, SD-JEPA improves over the LeWM baseline on the majority of its control benchmarks at matched compute, and outperforms the strongest non-LeWM JEPA baseline on Push-T; a subspace-ablation falsifier confirms the split is the load-bearing ingredient. Beyond planning, the resulting 1-D angular progression coordinate functions as a scene-aware compass on the latent. It advances with task progress, regresses when the agent backtracks, and under controlled perturbations both spikes and relocalises to a semantically appropriate new task-phase sector, separating the moment of surprise from its meaning in a way that prediction-error scalars cannot. Three quantitative tests back this up: $|Δθ_t|$ outperforms the standard latent-prediction-error surprise at localising semantic events on 40 held-out cube episodes by up to +0.18 pooled AUROC (97.5% per-episode win rate at $\pm 1$-step tolerance); a within-episode linear probe across all four environments (40 episodes per env) shows the 8-dimensional progression subspace (4.2% of the latent) explains 72-95% of task-progress variance..

📄 PDF Abstract BibTeX arXiv:2605.31111

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HARP: Hallucination Detection via Reasoning Subspace Projection

2025-09-15 · Junjie Hu, Gang Tu, ShengYu Cheng, Jinxin Li 외 arxiv

Hallucinations in Large Language Models (LLMs) pose a major barrier to their reliable use in critical decision-making. Although existing hallucination detection methods have improved accuracy, they still struggle with di…

Retriever: Learning Content-Style Representation as a Token-Level Bipartite Graph

2022-02-24 · ICLR 2022 4 · Dacheng Yin, Xuanchi Ren, Chong Luo, Yuwang Wang 외

This paper addresses the unsupervised learning of content-style decomposed representation. We first give a definition of style and then model the content-style representation as a token-level bipartite graph. An unsuperv…

DecoderQuantizationStyle TransferVoice Conversion

Exploiting video sequences for unsupervised disentangling in generative adversarial networks

2019-10-16 · Facundo Tuesca, Lucas C. Uzal

In this work we present an adversarial training algorithm that exploits correlations in video to learn --without supervision-- an image generator model with a disentangled latent space. The proposed methodology requires …

Koopman Invariants as Drivers of Emergent Time-Series Clustering in Joint-Embedding Predictive Architectures

2025-11-12 · Pablo Ruiz-Morales, Dries Vanoost, Davy Pissoort, Mathias Verbeke arxiv

Joint-Embedding Predictive Architectures (JEPAs), a powerful class of self-supervised models, exhibit an unexplained ability to cluster time-series data by their underlying dynamical regimes. We propose a novel theoretic…

Self-Supervised Learning

Illumination-Aware Age Progression

2014-06-01 · CVPR 2014 6 · Ira Kemelmacher-Shlizerman, Supasorn Suwajanakorn, Steven M. Seitz

We present an approach that takes a single photograph of a child as input and automatically produces a series of age-progressed outputs between 1 and 80 years of age, accounting for pose, expression, and illumination. Le…