paper-with-me

Papers

Temporally Centered SIGReg Improves LeWorldModel Representations for Robot Policy Learning

2026-07-29 · Chang Liu, Fei Suo, Yanzhou Jin, Yusuke Iwasawa, Yutaka Matsuo, Yaonan Zhu arxiv

Recent work on LeWorldModel (LeWM) has shown that the Sketched Isotropic Gaussian Regularizer (SIGReg) enables stable end-to-end world model learning from pixels by regularizing the latent representation toward an isotropic Gaussian. While effective for latent-space planning, the representations learned by Raw LeWM are poorly suited for downstream robot policy learning. In this paper, through Monte Carlo analysis, we show that the Raw LeWM objective biases variance allocation toward the temporally persistent component, thereby suppressing the variance of the temporally centered residual. Consistent with this analysis, trained Raw LeWM representations exhibit suppressed residual variation and reduced decodability of robot state and dynamics, particularly gripper dynamics, which are crucial for robotic manipulation. To address this issue, we apply SIGReg to temporally centered residuals rather than to the whole latent representation. This simple change decouples persistent and residual variance allocation while retaining an effective anti-collapse property. On the LIBERO benchmark, our method improves downstream policy success on the Goal suite by 1.66x and raises the average success rate across all suites from 63.6% to 83.8%. Without external pretraining, it also outperforms both Diffusion Policy trained from scratch and the pretrained OpenVLA baseline. These results associate the variance-allocation bias of Raw LeWM with the downstream policy gap, and show that decoupling persistent and residual variation yields representations better suited for downstream robot policy learning.

📄 PDF Abstract BibTeX arXiv:2607.26924

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models

2026-08-18 · Jack Boylan, Chris Hokamp arxiv

Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting future embeddings, but the objective admits a trivial solution of a constant encoder, so every practical system adds an anti-collapse mech…

SCALE: State-Calibrated Latent Embeddings for JEPA Planning in the Right Geometry

2026-08-17 · Jiaming Hu, Yan Zheng, Tian Wang arxiv

Joint-embedding predictive world models plan by scoring predicted terminal embeddings against a goal embedding using a cost defined on the representation itself. Two prominent strategies for obtaining non-collapsed repre…

Weak-SIGReg: Covariance Regularization for Stable Deep Learning

2026-03-06 · Habibullah Akbar arxiv

Modern neural network optimization relies heavily on architectural priorssuch as Batch Normalization and Residual connectionsto stabilize training dynamics. Without these, or in low-data regimes with aggressive augmentat…

Fast LeWorldModel

2026-06-24 · Yuntian Gao, Xiangyu Xu arxiv

Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising foundation for reconstruction-free visual world models. For visual planning, however, LeWM evaluates candida…

LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics

2025-11-11 · Randall Balestriero, Yann LeCun arxiv

Learning manipulable representations of the world and its dynamics is central to AI. Joint-Embedding Predictive Architectures (JEPAs) offer a promising blueprint, but lack of practical guidance and theory has led to ad-h…

Self-Supervised Learning