paper-with-me

홈 › Papers

Task Success is not Enough: Investigating the Use of Video-Language Models as Behavior Critics for Catching Undesirable Agent Behaviors

2024-02-06 · Lin Guan, Yifan Zhou, Denis Liu, Yantian Zha, Heni Ben Amor, Subbarao Kambhampati

Large-scale generative models are shown to be useful for sampling meaningful candidate solutions, yet they often overlook task constraints and user preferences. Their full power is better harnessed when the models are coupled with external verifiers and the final solutions are derived iteratively or progressively according to the verification feedback. In the context of embodied AI, verification often solely involves assessing whether goal conditions specified in the instructions have been met. Nonetheless, for these agents to be seamlessly integrated into daily life, it is crucial to account for a broader range of constraints and preferences beyond bare task success (e.g., a robot should grasp bread with care to avoid significant deformations). However, given the unbounded scope of robot tasks, it is infeasible to construct scripted verifiers akin to those used for explicit-knowledge tasks like the game of Go and theorem proving. This begs the question: when no sound verifier is available, can we use large vision and language models (VLMs), which are approximately omniscient, as scalable Behavior Critics to catch undesirable robot behaviors in videos? To answer this, we first construct a benchmark that contains diverse cases of goal-reaching yet undesirable robot policies. Then, we comprehensively evaluate VLM critics to gain a deeper understanding of their strengths and failure modes. Based on the evaluation, we provide guidelines on how to effectively utilize VLM critiques and showcase a practical way to integrate the feedback into an iterative process of policy refinement. The dataset and codebase are released at: https://guansuns.github.io/pages/vlm-critic.

📄 PDF Abstract BibTeX arXiv:2402.04210

Code (0)

등록된 구현이 없습니다.

Tasks

Automated Theorem ProvingGame of Go

Similar Papers 제목 키워드 기반

Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding

2024-03-14 · Guo Chen, Yifei HUANG, Jilan Xu, Baoqi Pei 외

Understanding videos is one of the fundamental directions in computer vision research, with extensive efforts dedicated to exploring various architectures such as RNN, 3D CNN, and Transformers. The newly proposed archite…

MambaMoment RetrievalTemporal Action LocalizationVideo Understanding

Distilling Vision-Language Models on Millions of Videos

2024-01-11 · CVPR 2024 1 · Yue Zhao, Long Zhao, Xingyi Zhou, Jialin Wu 외

The recent advance in vision-language models is largely attributed to the abundance of image-text data. We aim to replicate this success for video-language models, but there simply is not enough human-curated video-text …

Language ModelingLanguage ModellingRetrievalText to Video Retrieval+1

A Rough Set Formalization of Quantitative Evaluation with Ambiguity

2012-05-01 · LREC 2012 5 · Patrick Paroubek, Xavier Tannier

In this paper, we present the founding elements of a formal model of the evaluation paradigm in natural language processing. We propose an abstract model of objective quantitative evaluation based on rough sets, as well …

Information RetrievalMachine TranslationNamed Entity Recognition (NER)Word Sense Disambiguation

Live Repetition Counting

2015-12-01 · ICCV 2015 12 · Ofir Levy, Lior Wolf

The task of counting the number of repetitions of approximately the same action in an input video sequence is addressed. The proposed method runs online and not on the complete pre-captured video. It analyzes sequentiall…

Can Everybody Sign Now? Exploring Sign Language Video Generation from 2D Poses

2020-12-20 · Lucas Ventura, Amanda Duarte, Xavier Giro-i-Nieto

Recent work have addressed the generation of human poses represented by 2D/3D coordinates of human joints for sign language. We use the state of the art in Deep Learning for motion transfer and evaluate them on How2Sign,…

Sign Language ProductionVideo Generation