paper-with-me

홈 › Papers

Latent Space Roadmap for Visual Action Planning of Deformable and Rigid Object Manipulation

2020-03-19 · Martina Lippi, Petra Poklukar, Michael C. Welle, Anastasiia Varava, Hang Yin, Alessandro Marino, Danica Kragic

We present a framework for visual action planning of complex manipulation tasks with high-dimensional state spaces such as manipulation of deformable objects. Planning is performed in a low-dimensional latent state space that embeds images. We define and implement a Latent Space Roadmap (LSR) which is a graph-based structure that globally captures the latent system dynamics. Our framework consists of two main components: a Visual Foresight Module (VFM) that generates a visual plan as a sequence of images, and an Action Proposal Network (APN) that predicts the actions between them. We show the effectiveness of the method on a simulated box stacking task as well as a T-shirt folding task performed with a real robot.

📄 PDF Abstract BibTeX arXiv:2003.08974

Code (1)

visual-action-planning/lsr-code 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Enabling Visual Action Planning for Object Manipulation through Latent Space Roadmap

2021-03-03 · Martina Lippi, Petra Poklukar, Michael C. Welle, Anastasia Varava 외

We present a framework for visual action planning of complex manipulation tasks with high-dimensional state spaces, focusing on manipulation of deformable objects. We propose a Latent Space Roadmap (LSR) for task plannin…

Task Planning

CTRMs: Learning to Construct Cooperative Timed Roadmaps for Multi-agent Path Planning in Continuous Spaces

2022-01-24 · Keisuke Okumura, Ryo Yonetani, Mai Nishimura, Asako Kanezaki

Multi-agent path planning (MAPP) in continuous spaces is a challenging problem with significant practical importance. One promising approach is to first construct graphs approximating the spaces, called roadmaps, and the…

CLAD: Constrained Latent Action Diffusion for Vision-Language Procedure Planning

2025-03-09 · Lei Shi, Andreas Bulling

We propose CLAD -- a Constrained Latent Action Diffusion model for vision-language procedure planning in instructional videos. Procedure planning is the challenging task of predicting intermediate actions given a visual …

Simulating the Visual World with Artificial Intelligence: A Roadmap

2025-11-11 · Jingtong Yue, Ziqi Huang, Zhaoxi Chen, Xintao Wang 외 arxiv

The landscape of video generation is shifting, from a focus on generating visually appealing clips to building virtual environments that support interaction and maintain physical plausibility. These developments point to…

Autonomous DrivingVisual ReasoningVideo Generation

Trajectory-Constrained Deep Latent Visual Attention for Improved Local Planning in Presence of Heterogeneous Terrain

2021-12-09 · Stefan Wapnick, Travis Manderson, David Meger, Gregory Dudek

We present a reward-predictive, model-based deep learning method featuring trajectory-constrained visual attention for local planning in visual navigation tasks. Our method learns to place visual attention at locations i…

Visual Navigation