paper-with-me

홈 › Papers

VOILA: Visual-Observation-Only Imitation Learning for Autonomous Navigation

2021-05-19 · Haresh Karnan, Garrett Warnell, Xuesu Xiao, Peter Stone

While imitation learning for vision based autonomous mobile robot navigation has recently received a great deal of attention in the research community, existing approaches typically require state action demonstrations that were gathered using the deployment platform. However, what if one cannot easily outfit their platform to record these demonstration signals or worse yet the demonstrator does not have access to the platform at all? Is imitation learning for vision based autonomous navigation even possible in such scenarios? In this work, we hypothesize that the answer is yes and that recent ideas from the Imitation from Observation (IfO) literature can be brought to bear such that a robot can learn to navigate using only ego centric video collected by a demonstrator, even in the presence of viewpoint mismatch. To this end, we introduce a new algorithm, Visual Observation only Imitation Learning for Autonomous navigation (VOILA), that can successfully learn navigation policies from a single video demonstration collected from a physically different agent. We evaluate VOILA in the photorealistic AirSim simulator and show that VOILA not only successfully imitates the expert, but that it also learns navigation policies that can generalize to novel environments. Further, we demonstrate the effectiveness of VOILA in a real world setting by showing that it allows a wheeled Jackal robot to successfully imitate a human walking in an environment using a video recorded using a mobile phone camera.

📄 PDF Abstract BibTeX arXiv:2105.09371

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous NavigationImitation LearningNavigateRobot Navigation

Similar Papers 제목 키워드 기반

VOiLA: Vectorized Online Planning with Learned Diffusion Models for POMDP Agents

2026-06-18 · Marcus Hoerger, Rishikesh Joshi, Rahul Shome, Ian Manchester 외 arxiv

Planning under uncertainty is an essential capability for autonomous robots. The Partially Observable Markov Decision Process (POMDP) provides a powerful framework for such a capability. Although POMDP-based planning has…

VOILA: An Optimised Dialogue System for Interactively Learning Visually-Grounded Word Meanings (Demonstration System)

2017-08-01 · WS 2017 8 · Yanchao Yu, Arash Eshghi, Oliver Lemon

We present VOILA: an optimised, multi-modal dialogue agent for interactive learning of visually grounded word meanings from a human user. VOILA is: (1) able to learn new visual categories interactively from users from sc…

Active Learning

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

2025-05-05 · Yemin Shi, Yu Shu, Siwei Dong, Guangyi Liu 외

A voice AI agent that blends seamlessly into daily life would interact with humans in an autonomous, real-time, and emotionally expressive manner. Rather than merely reacting to commands, it would continuously listen, re…

AI AgentAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Rhythm+4

VOILA: Value-of-Information Guided Fidelity Selection for Cost-Aware Multimodal Question Answering

2026-02-03 · Rahul Atul Bhope, K. R. Jayaram, Vinod Muthusamy, Ritesh Kumar 외 arxiv

Despite significant costs from retrieving and processing high-fidelity visual inputs, most multimodal vision-language systems operate at fixed fidelity levels. We introduce VOILA, a framework for Value-Of-Information-dri…

Visual Question Answering

G-VOILA: Gaze-Facilitated Information Querying in Daily Scenarios

2024-05-13 · Zeyu Wang, Yuanchun Shi, Yuntao Wang, Yuchen Yao 외

Modern information querying systems are progressively incorporating multimodal inputs like vision and audio. However, the integration of gaze -- a modality deeply linked to user intent and increasingly accessible via gaz…

Natural Language Queries