paper-with-me

홈 › Papers

LIBERO-VIFO: Benchmarking the Capability and Safety of Visual Cue Following in Vision-Language-Action Models

2026-08-18 · Zhengyan Qian, Rui Yan, Alex Jinpeng Wang, Jinhui Tang arxiv

Visual cues are increasingly adopted to guide robot learning, but whether Vision-Language-Action (VLA) models can reliably follow authorized cues while disregarding unauthorized ones remains unclear. Existing work covers only a narrow range of cue forms and focuses on final task success, providing only a coarse assessment of cue-following capability. Treating all visual cues as authorized also leaves safety risks of unauthorized following unexplored. To address these gaps, we introduce LIBERO-VIFO, a benchmark to evaluate both the capability and safety of visual cue following in VLA models. LIBERO-VIFO defines eight visual cue families spanning diverse forms. A total of four protocols in two parts are defined: Part I tests cue understanding and authorized following, while Part II evaluates unauthorized visual cue following under language-cue conflict and empty language conditions. Evaluating seven VLA models reveals that although visual cue understanding does not reliably translate into execution, current VLAs are able to execute cue-indicated tasks without language instruction, exposing an emerging risk of unauthorized visual cue following. Extended experiments on scene-instantiated cues, safety-critical settings, and real-robot deployment corroborate these findings. LIBERO-VIFO brings both the capability and safety of visual cue following into systematic evaluation, establishing visual-centric safety as a new perspective for the VLA community.

📄 PDF Abstract BibTeX arXiv:2608.17600

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

What Are We Actually Benchmarking in Robot Manipulation?

2026-06-02 · Tianchong Jiang, Xiangshan Tan, Samuel Wheeler, Luzhe Sun 외 arxiv

A robotics benchmark score measures success under one fixed evaluation setup, yet is routinely treated as evidence of general manipulation capability. We identify four failure modes, each of which weakens or invalidates …

Robot Manipulation

Variational Inference on the Final-Layer Output of Neural Networks

2023-02-05 · Yadi Wei, Roni Khardon

Traditional neural networks are simple to train but they typically produce overconfident predictions. In contrast, Bayesian neural networks provide good uncertainty quantification but optimizing them is time consuming du…

Uncertainty QuantificationVariational Inference

NaviFormer: A Deep Reinforcement Learning Transformer-like Model to Holistically Solve the Navigation Problem

2026-04-18 · Daniel Fuertes, Andrea Cavallaro, Carlos R. del-Blanco, Fernando Jaureguizar 외 arxiv

Path planning is usually solved by addressing either the (high-level) route planning problem (waypoint sequencing to achieve the final goal) or the (low-level) path planning problem (trajectory prediction between two way…

Reinforcement LearningTrajectory Prediction

LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

2023-06-05 · NeurIPS 2023 11

Lifelong learning offers a promising paradigm of building a generalist agent that learns and adapts over its lifespan. Unlike traditional lifelong learning problems in image and text domains, which primarily involve the …

BenchmarkingDecision MakingLifelong learning+2

RealViformer: Investigating Attention for Real-World Video Super-Resolution

2024-07-19 · Yuehan Zhang, Angela Yao

In real-world video super-resolution (VSR), videos suffer from in-the-wild degradations and artifacts. VSR methods, especially recurrent ones, tend to propagate artifacts over time in the real-world setting and are more …

Image Super-ResolutionSuper-ResolutionVideo Super-Resolution