paper-with-me

홈 › Papers

Grounding Video Models to Actions through Goal Conditioned Exploration

2024-11-11 · Yunhao Luo, Yilun Du

Large video models, pretrained on massive amounts of Internet video, provide a rich source of physical knowledge about the dynamics and motions of objects and tasks. However, video models are not grounded in the embodiment of an agent, and do not describe how to actuate the world to reach the visual states depicted in a video. To tackle this problem, current methods use a separate vision-based inverse dynamic model trained on embodiment-specific data to map image states to actions. Gathering data to train such a model is often expensive and challenging, and this model is limited to visual settings similar to the ones in which data are available. In this paper, we investigate how to directly ground video models to continuous actions through self-exploration in the embodied environment -- using generated video states as visual goals for exploration. We propose a framework that uses trajectory level action generation in combination with video guidance to enable an agent to solve complex tasks without any external supervision, e.g., rewards, action labels, or segmentation masks. We validate the proposed approach on 8 tasks in Libero, 6 tasks in MetaWorld, 4 tasks in Calvin, and 12 tasks in iThor Visual Navigation. We show how our approach is on par with or even surpasses multiple behavior cloning baselines trained on expert demonstrations while without requiring any action annotations.

📄 PDF Abstract BibTeX arXiv:2411.07223

Code (0)

등록된 구현이 없습니다.

Tasks

Action GenerationVisual Navigation

Similar Papers 제목 키워드 기반

Grounding Generated Videos in Feasible Plans via World Models

2026-02-02 · Christos Ziakas, Amir Bar, Alessandra Russo arxiv

Large-scale video generative models have shown emerging capabilities as zero-shot visual planners, yet video-generated plans often violate temporal consistency and physical constraints, leading to failures when mapped to…

Translating Flow to Policy via Hindsight Online Imitation

2025-12-22 · Yitian Zheng, Zhangchen Ye, Weijun Dong, Shengjie Wang 외 arxiv

Recent advances in hierarchical robot systems leverage a high-level planner to propose task plans and a low-level policy to generate robot actions. This design allows training the planner on action-free or even non-robot…

Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals

2026-01-09 · Nate Gillman, Yinghua Zhou, Zitian Tang, Evan Luo 외 arxiv

Recent advancements in video generation have enabled the development of ``world models'' capable of simulating potential futures for robotics and planning. However, specifying precise goals for these models remains a cha…

Zero-shot GeneralizationVideo Generation

Language-Conditioned Goal Generation: a New Approach to Language Grounding for RL

2020-06-12 · Cédric Colas, Ahmed Akakzia, Pierre-Yves Oudeyer, Mohamed Chetouani 외

In the real world, linguistic agents are also embodied agents: they perceive and act in the physical world. The notion of Language Grounding questions the interactions between language and embodiment: how do learning age…

DiversityInstruction FollowingLanguage Acquisition

Language-Conditioned Goal Generation: a New Approach to Language Grounding in RL

2020-06-12 · ICML Workshop LaReL 2020 7 · Cédric Colas, Ahmed Akakzia, Pierre-Yves Oudeyer, Mohamed Chetouani 외

In the real world, linguistic agents are also embodied agents: they perceive and act in the physical world. The notion of Language Grounding questions the interactions between language and embodiment: how do learning age…

DiversityInstruction FollowingLanguage Acquisition