paper-with-me

홈 › Papers

Learning Visual Locomotion with Cross-Modal Supervision

2022-11-07 · Antonio Loquercio, Ashish Kumar, Jitendra Malik

In this work, we show how to learn a visual walking policy that only uses a monocular RGB camera and proprioception. Since simulating RGB is hard, we necessarily have to learn vision in the real world. We start with a blind walking policy trained in simulation. This policy can traverse some terrains in the real world but often struggles since it lacks knowledge of the upcoming geometry. This can be resolved with the use of vision. We train a visual module in the real world to predict the upcoming terrain with our proposed algorithm Cross-Modal Supervision (CMS). CMS uses time-shifted proprioception to supervise vision and allows the policy to continually improve with more real-world experience. We evaluate our vision-based walking policy over a diverse set of terrains including stairs (up to 19cm high), slippery slopes (inclination of 35 degrees), curbs and tall steps (up to 20cm), and complex discrete terrains. We achieve this performance with less than 30 minutes of real-world data. Finally, we show that our policy can adapt to shifts in the visual field with a limited amount of real-world experience. Video results and code at https://antonilo.github.io/vision_locomotion/.

📄 PDF Abstract BibTeX arXiv:2211.03785

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RevalExo: A Functional Daily-Activity Benchmark for Inertial and Visual Locomotion Mode Recognition in Older Adults and Clinical Cohorts

2026-09-08 · Diwas Lamsal, Juha Carlon, Reinhard Claeys, Maxim Yudayev 외 arxiv

Assistive devices for people with mobility impairments, such as powered exoskeletons, rely on accurate locomotion mode recognition to adapt control strategies and provide appropriate assistance during daily activities. H…

WOLF-VLA: Whole-Body Humanoid Optimal Locomotion Framework for Vision-Language-Action Learning

2026-06-24 · Melya Boukheddimi, Omar Adjali, Daniel Sontag, Frank Kirchner arxiv

Vision-Language-Action (VLA) models have recently demonstrated strong generalization in robotic manipulation, yet their applicability to whole-body, contact-rich humanoid locomotion remains severely underexplored due to …

Motion Synthesis

CReF: Cross-modal and Recurrent Fusion for Depth-conditioned Humanoid Locomotion

2026-03-31 · Yuan Hao, Ruiqi Yu, Shixin Luo, Guoteng Zhang 외 arxiv

Stable traversal over geometrically complex terrain increasingly requires exteroceptive perception, yet prior perceptive humanoid locomotion methods often remain tied to explicit geometric abstractions, either by mediati…

MCRL4OR: Multimodal Contrastive Representation Learning for Off-Road Environmental Perception

2025-01-23 · Yi Yang, Zhang Zhang, Liang Wang

Most studies on environmental perception for autonomous vehicles (AVs) focus on urban traffic environments, where the objects/stuff to be perceived are mainly from man-made scenes and scalable datasets with dense annotat…

Autonomous VehiclesContrastive LearningRepresentation Learning

Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing

2025-09-17 · Yaru Chen, Ruohao Guo, Liting Gao, Yang Xiang 외 arxiv

Weakly-supervised audio-visual video parsing (AVVP) seeks to detect audible, visible, and audio-visual events without temporal annotations. Previous work has emphasized refining global predictions through contrastive or …