paper-with-me

홈 › Papers

Commonsense Reasoning for Legged Robot Adaptation with Vision-Language Models

2024-07-02 · Annie S. Chen, Alec M. Lessing, Andy Tang, Govind Chada, Laura Smith, Sergey Levine, Chelsea Finn

Legged robots are physically capable of navigating a diverse variety of environments and overcoming a wide range of obstructions. For example, in a search and rescue mission, a legged robot could climb over debris, crawl through gaps, and navigate out of dead ends. However, the robot's controller needs to respond intelligently to such varied obstacles, and this requires handling unexpected and unusual scenarios successfully. This presents an open challenge to current learning methods, which often struggle with generalization to the long tail of unexpected situations without heavy human supervision. To address this issue, we investigate how to leverage the broad knowledge about the structure of the world and commonsense reasoning capabilities of vision-language models (VLMs) to aid legged robots in handling difficult, ambiguous situations. We propose a system, VLM-Predictive Control (VLM-PC), combining two key components that we find to be crucial for eliciting on-the-fly, adaptive behavior selection with VLMs: (1) in-context adaptation over previous robot interactions and (2) planning multiple skills into the future and replanning. We evaluate VLM-PC on several challenging real-world obstacle courses, involving dead ends and climbing and crawling, on a Go1 quadruped robot. Our experiments show that by reasoning over the history of interactions and future plans, VLMs enable the robot to autonomously perceive, navigate, and act in a wide range of complex scenarios that would otherwise require environment-specific engineering or human guidance.

📄 PDF Abstract BibTeX arXiv:2407.02666

Code (0)

등록된 구현이 없습니다.

Tasks

Navigate

Similar Papers 제목 키워드 기반

LocoVLM: Grounding Vision and Language for Adapting Versatile Legged Locomotion Policies

2026-02-11 · I Made Aswin Nahrendra, Seunghyun Lee, Dongkyu Lee, Hyun Myung arxiv

Recent advances in legged locomotion learning are still dominated by the utilization of geometric representations of the environment, limiting the robot's capability to respond to higher-level semantics such as human ins…

RMA: Rapid Motor Adaptation for Legged Robots

2021-07-08 · Ashish Kumar, Zipeng Fu, Deepak Pathak, Jitendra Malik

Successful real-world deployment of legged robots would require them to adapt in real-time to unseen scenarios like changing terrains, changing payloads, wear and tear. This paper presents Rapid Motor Adaptation (RMA) al…

Sand

Toward Grounded Commonsense Reasoning

2023-06-14 · Minae Kwon, Hengyuan Hu, Vivek Myers, Siddharth Karamcheti 외

Consider a robot tasked with tidying a desk with a meticulously constructed Lego sports car. A human may recognize that it is not appropriate to disassemble the sports car and put it away as part of the "tidying." How ca…

Language Modelling

PhysBrain 1.0 Technical Report

2026-05-14 · Shijie Lian, Bin Yu, Xiaopeng Lin, Changti Wu 외 arxiv

Vision-language-action models have advanced rapidly, but robot trajectories alone provide limited coverage for learning broad physical understanding. PhysBrain 1.0 studies a complementary route: converting large-scale hu…

Enabling Robots to Understand Incomplete Natural Language Instructions Using Commonsense Reasoning

2019-04-29 · Haonan Chen, Hao Tan, Alan Kuntz, Mohit Bansal 외

Enabling robots to understand instructions provided via spoken natural language would facilitate interaction between robots and people in a variety of settings in homes and workplaces. However, natural language instructi…

Common Sense ReasoningLanguage ModelingLanguage Modelling