paper-with-me

Papers

MVP: Unified Motion and Visual Self-Supervised Learning for Large-Scale Robotic Navigation

2020-03-02 · Marvin Chancán, Michael Milford

Autonomous navigation emerges from both motion and local visual perception in real-world environments. However, most successful robotic motion estimation methods (e.g. VO, SLAM, SfM) and vision systems (e.g. CNN, visual place recognition-VPR) are often separately used for mapping and localization tasks. Conversely, recent reinforcement learning (RL) based methods for visual navigation rely on the quality of GPS data reception, which may not be reliable when directly using it as ground truth across multiple, month-spaced traversals in large environments. In this paper, we propose a novel motion and visual perception approach, dubbed MVP, that unifies these two sensor modalities for large-scale, target-driven navigation tasks. Our MVP-based method can learn faster, and is more accurate and robust to both extreme environmental changes and poor GPS data than corresponding vision-only navigation methods. MVP temporally incorporates compact image representations, obtained using VPR, with optimized motion estimation data, including but not limited to those from VO or optimized radar odometry (RO), to efficiently learn self-supervised navigation policies via RL. We evaluate our method on two large real-world datasets, Oxford Robotcar and Nordland Railway, over a range of weather (e.g. overcast, night, snow, sun, rain, clouds) and seasonal (e.g. winter, spring, fall, summer) conditions using the new CityLearn framework; an interactive environment for efficiently training navigation agents. Our experimental results, on traversals of the Oxford RobotCar dataset with no GPS data, show that MVP can achieve 53% and 93% navigation success rate using VO and RO, respectively, compared to 7% for a vision-only method. We additionally report a trade-off between the RL success rate and the motion estimation precision.

📄 PDF Abstract BibTeX arXiv:2003.00667

Code (1)

mchancan/citylearn 공식 구현

Tasks

Autonomous DrivingAutonomous NavigationAutonomous VehiclesMotion EstimationRadar odometryReinforcement LearningReinforcement Learning (RL)Robot NavigationSelf-Driving CarsSelf-Supervised LearningVisual LocalizationVisual NavigationVisual Place Recognition

Similar Papers 제목 키워드 기반

SyncLipMAE: Contrastive Masked Pretraining for Audio-Visual Talking-Face Representation

2025-10-11 · Zeyu Ling, Xiaodong Gu, Jiangnan Tang, Changqing Zou arxiv

We introduce SyncLipMAE, a self-supervised pretraining framework for talking-face video that learns synchronization-aware and transferable facial dynamics from unlabeled audio-visual streams. Our approach couples masked …

Visual Speech RecognitionAction Recognition

HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition

2024-01-11 · Licai Sun, Zheng Lian, Bin Liu, JianHua Tao

Audio-Visual Emotion Recognition (AVER) has garnered increasing attention in recent years for its critical role in creating emotion-ware intelligent machines. Previous efforts in this area are dominated by the supervised…

Contrastive LearningDynamic Facial Expression RecognitionEmotion RecognitionRepresentation Learning+1

SSAVSV: Towards Unified Model for Self-Supervised Audio-Visual Speaker Verification

2025-06-21 · Gnana Praveen Rajasekhar, Jahangir Alam

Conventional audio-visual methods for speaker verification rely on large amounts of labeled data and separate modality-specific architectures, which is computationally expensive, limiting their scalability. To address th…

Contrastive LearningSelf-Supervised LearningSpeaker Verification

VAEmo: Efficient Representation Learning for Visual-Audio Emotion with Knowledge Injection

2025-05-05 · Hao Cheng, Zhiwei Zhao, Yichao He, Zhenzhen Hu 외

Audiovisual emotion recognition (AVER) aims to infer human emotions from nonverbal visual-audio (VA) cues, offering modality-complementary and language-agnostic advantages. However, AVER remains challenging due to the in…

Contrastive LearningDynamic Facial Expression RecognitionEmotion RecognitionRepresentation Learning+2

Self-Supervised Monocular Depth and Ego-Motion Estimation in Endoscopy: Appearance Flow to the Rescue

2021-12-15 · Shuwei Shao, Zhongcai Pei, Weihai Chen, Wentao Zhu 외

Recently, self-supervised learning technology has been applied to calculate depth and ego-motion from monocular videos, achieving remarkable performance in autonomous driving scenarios. One widely adopted assumption of d…

Depth EstimationMotion EstimationSelf-Supervised Learning