paper-with-me

홈 › Papers

Compositional Foundation Models for Hierarchical Planning

2023-09-15 · NeurIPS 2023 11

To make effective decisions in novel environments with long-horizon goals, it is crucial to engage in hierarchical reasoning across spatial and temporal scales. This entails planning abstract subgoal sequences, visually reasoning about the underlying plans, and executing actions in accordance with the devised plan through visual-motor control. We propose Compositional Foundation Models for Hierarchical Planning (HiP), a foundation model which leverages multiple expert foundation model trained on language, vision and action data individually jointly together to solve long-horizon tasks. We use a large language model to construct symbolic plans that are grounded in the environment through a large video diffusion model. Generated video plans are then grounded to visual-motor control, through an inverse dynamics model that infers actions from generated videos. To enable effective reasoning within this hierarchy, we enforce consistency between the models via iterative refinement. We illustrate the efficacy and adaptability of our approach in three different long-horizon table-top manipulation tasks.

📄 PDF Abstract BibTeX arXiv:2309.08587

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language Model

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Compositional planning in Markov decision processes: Temporal abstraction meets generalized logic composition

2018-10-05 · Xuan Liu, Jie Fu

In hierarchical planning for Markov decision processes (MDPs), temporal abstraction allows planning with macro-actions that take place at different time scale in form of sequential composition. In this paper, we propose …

LVLM-Composer's Explicit Planning for Image Generation

2025-07-05 · Spencer Ramsey, Jeffrey Lee, Amina Grant arxiv

The burgeoning field of generative artificial intelligence has fundamentally reshaped our approach to content creation, with Large Vision-Language Models (LVLMs) standing at its forefront. While current LVLMs have demons…

Text-to-Image GenerationReinforcement LearningVisual Grounding

Simple Hierarchical Planning with Diffusion

2024-01-05 · Chang Chen, Fei Deng, Kenji Kawaguchi, Caglar Gulcehre 외

Diffusion-based generative methods have proven effective in modeling trajectories with offline datasets. However, they often face computational challenges and can falter in generalization, especially in capturing tempora…

NEUSIS: A Compositional Neuro-Symbolic Framework for Autonomous Perception, Reasoning, and Planning in Complex UAV Search Missions

2024-09-16 · Zhixi Cai, Cristian Rojas Cardenas, Kevin Leo, Chenyuan Zhang 외

This paper addresses the problem of autonomous UAV search missions, where a UAV must locate specific Entities of Interest (EOIs) within a time limit, based on brief descriptions in large, hazard-prone environments with k…

Language ModelingLanguage Modelling

RoboHiMan: A Hierarchical Evaluation Paradigm for Compositional Generalization in Long-Horizon Manipulation

2025-10-15 · Yangtao Chen, Zixuan Chen, Nga Teng Chan, Junting Chen 외 arxiv

Enabling robots to flexibly schedule and compose learned skills for novel long-horizon manipulation under diverse perturbations remains a core challenge. Early explorations with end-to-end VLA models show limited success…