paper-with-me

홈 › Papers

V2CNet: A Deep Learning Framework to Translate Videos to Commands for Robotic Manipulation

2019-03-23 · Anh Nguyen, Thanh-Toan Do, Ian Reid, Darwin G. Caldwell, Nikos G. Tsagarakis

We propose V2CNet, a new deep learning framework to automatically translate the demonstration videos to commands that can be directly used in robotic applications. Our V2CNet has two branches and aims at understanding the demonstration video in a fine-grained manner. The first branch has the encoder-decoder architecture to encode the visual features and sequentially generate the output words as a command, while the second branch uses a Temporal Convolutional Network (TCN) to learn the fine-grained actions. By jointly training both branches, the network is able to model the sequential information of the command, while effectively encodes the fine-grained actions. The experimental results on our new large-scale dataset show that V2CNet outperforms recent state-of-the-art methods by a substantial margin, while its output can be applied in real robotic applications. The source code and trained models will be made available.

📄 PDF Abstract BibTeX arXiv:1903.10869

Code (0)

등록된 구현이 없습니다.

Tasks

Decoder

Similar Papers 제목 키워드 기반

Translating Videos to Commands for Robotic Manipulation with Deep Recurrent Neural Networks

2017-10-01 · Anh Nguyen, Dimitrios Kanoulas, Luca Muratore, Darwin G. Caldwell 외

We present a new method to translate videos to commands for robotic manipulation using Deep Recurrent Neural Networks (RNN). Our framework first extracts deep features from the input video frames with a deep Convolutiona…

DecoderTranslation

Learning Lexical Entries for Robotic Commands using Crowdsourcing

2016-09-08 · Junjie Hu, Jean Oh, Anatole Gershman

Robotic commands in natural language usually contain various spatial descriptions that are semantically similar but syntactically different. Mapping such syntactic variants into semantic concepts that can be understood b…

Machine TranslationTranslation

Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow

2025-12-31 · Karthik Dharmarajan, Wenlong Huang, Jiajun Wu, Li Fei-Fei 외 arxiv

Generative video modeling has emerged as a compelling tool to zero-shot reason about plausible physical interactions for open-world manipulation. Yet, it remains a challenge to translate such human-led motions into the l…

Reinforcement LearningVideo Generation

How Does it Sound?

2021-12-01 · NeurIPS 2021 12 · Kun Su, Xiulong Liu, Eli Shlizerman

One of the primary purposes of video is to capture people and their unique activities. It is often the case that the experience of watching the video can be enhanced by adding a musical soundtrack that is in-sync with th…

Rhythm

Robotic Grasping and Placement Controlled by EEG-Based Hybrid Visual and Motor Imagery

2026-03-03 · Yichang Liu, Tianyu Wang, Ziyi Ye, Yawei Li 외 arxiv

We present a framework that integrates EEG-based visual and motor imagery (VI/MI) with robotic control to enable real-time, intention-driven grasping and placement. Motivated by the promise of BCI-driven robotics to enha…

Robotic Grasping