paper-with-me

홈 › Papers

Back to Parsimonious Latents: Learning Task-Centric World Models from Visual Foundations

2026-05-25 · Minghao Fu, Fan Feng, Nicklas Hansen, Biwei Huang arxiv

World models enable agents to predict future dynamics conditioned on actions, making the choice of latent representation central to planning and control. Such representations are often either learned directly from pixels with limited semantic structure or inherited from frozen visual foundation models with excessive task-irrelevant detail, yielding state spaces that are poorly matched to downstream planning and control. This is especially challenging in reward-free offline settings, where the model must learn from fixed trajectories without reward supervision or online interaction. To address this, we propose TC-WM, a framework for turning foundation-model embeddings into compact, task-sufficient world representations. The key design is to treat the pretrained embedding space as a semantic scaffold rather than as the final state space: TC-WM linearly projects high-dimensional visual embeddings into a compact latent as the dynamic space, aligns a subspace with the agent's physical state via contrastive learning, and reconstructs embeddings to preserve useful visual structure. This combines the generality of foundation features with the controllability of task-centric dynamics. Theoretically, we show that TC-WM suffices to identify the underlying task-centric latent factors up to a simple transformation. Empirically, TC-WM enables test-time planning across diverse environments (e.g., Robomimic and D4RL), achieving better world-modeling quality and more precise control than state-of-the-art approaches.

📄 PDF Abstract BibTeX arXiv:2605.25620

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive Learning

Similar Papers 제목 키워드 기반

$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

2026-08-06 · Zhe Li, Zhenzhe Zhang, Yangyang Wei, Wenjie Zhang 외 arxiv

Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies ty…

When Object-Centric World Models Meet Policy Learning: From Pixels to Policies, and Where It Breaks

2025-11-08 · Stefano Ferraro, Akihiro Nakano, Masahiro Suzuki, Yutaka Matsuo arxiv

Object-centric world models (OCWM) aim to decompose visual scenes into object-level representations, providing structured abstractions that could improve compositional generalization and data efficiency in reinforcement …

Reinforcement Learning

HiMem-WAM: Hierarchical Memory-Gated World Action Models for Robotic Manipulation

2026-06-09 · Xiaoquan Sun, Ruijian Zhang, Chen Cao, Yihan Sun 외 arxiv

World Action Models (WAMs) have emerged as a new powerful paradigm for embodied intelligence, learning action-relevant visual dynamics that significantly enhance generalization and robustness. However, existing WAMs stil…

Optical Flow EstimationCausal Inference

FlexEdit: Flexible and Controllable Diffusion-based Object-centric Image Editing

2024-03-27 · Trong-Tung Nguyen, Duc-Anh Nguyen, Anh Tran, Cuong Pham

Our work addresses limitations seen in previous approaches for object-centric editing problems, such as unrealistic results due to shape discrepancies and limited control in object replacement or insertion. To this end, …

DenoisingObjecttext-guided-image-editing

SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos

2022-06-15 · Gamaleldin F. Elsayed, Aravindh Mahendran, Sjoerd van Steenkiste, Klaus Greff 외

The visual world can be parsimoniously characterized in terms of distinct entities with sparse interactions. Discovering this compositional structure in dynamic visual scenes has proven challenging for end-to-end compute…

ObjectSemantic Segmentation