paper-with-me

홈 › Papers

Translating Videos to Commands for Robotic Manipulation with Deep Recurrent Neural Networks

2017-10-01 · Anh Nguyen, Dimitrios Kanoulas, Luca Muratore, Darwin G. Caldwell, Nikos G. Tsagarakis

We present a new method to translate videos to commands for robotic manipulation using Deep Recurrent Neural Networks (RNN). Our framework first extracts deep features from the input video frames with a deep Convolutional Neural Networks (CNN). Two RNN layers with an encoder-decoder architecture are then used to encode the visual features and sequentially generate the output words as the command. We demonstrate that the translation accuracy can be improved by allowing a smooth transaction between two RNN layers and using the state-of-the-art feature extractor. The experimental results on our new challenging dataset show that our approach outperforms recent methods by a fair margin. Furthermore, we combine the proposed translation module with the vision and planning system to let a robot perform various manipulation tasks. Finally, we demonstrate the effectiveness of our framework on a full-size humanoid robot WALK-MAN.

📄 PDF Abstract BibTeX arXiv:1710.00290

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderTranslation

Similar Papers 제목 키워드 기반

Learning Lexical Entries for Robotic Commands using Crowdsourcing

2016-09-08 · Junjie Hu, Jean Oh, Anatole Gershman

Robotic commands in natural language usually contain various spatial descriptions that are semantically similar but syntactically different. Mapping such syntactic variants into semantic concepts that can be understood b…

Machine TranslationTranslation

Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation

2025-08-30 · Chuye Zhang, Xiaoxiong Zhang, Wei Pan, Linfang Zheng 외 arxiv

Robotic manipulation in unstructured environments requires systems that can generalize across diverse tasks while maintaining robust and reliable performance. We introduce {GVF-TAPE}, a closed-loop framework that combine…

Pose Estimation

V2CNet: A Deep Learning Framework to Translate Videos to Commands for Robotic Manipulation

2019-03-23 · Anh Nguyen, Thanh-Toan Do, Ian Reid, Darwin G. Caldwell 외

We propose V2CNet, a new deep learning framework to automatically translate the demonstration videos to commands that can be directly used in robotic applications. Our V2CNet has two branches and aims at understanding th…

Decoder

Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills

2025-03-16 · Haoqi Yuan, Yu Bai, Yuhui Fu, Bohan Zhou 외

Building autonomous robotic agents capable of achieving human-level performance in real-world embodied tasks is an ultimate goal in humanoid robot research. Recent advances have made significant progress in high-level co…

Task Planning

Robot Learning from a Physical World Model

2025-11-10 · Jiageng Mao, Sicheng He, Hao-Ning Wu, Yang You 외 arxiv

We introduce PhysWorld, a framework that enables robot learning from video generation through physical world modeling. Recent video generation models can synthesize photorealistic visual demonstrations from language comm…

Reinforcement LearningVideo Generation