paper-with-me

홈 › Papers

UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models

2026-02-20 · Jiabing Yang, Yixiang Chen, Yuan Xu, Peiyan Li, Zichen Wen, Bowen Fang, Tao Yu, Xiangnan Wu, Qisen Ma, Kai Wang, Ziheng He, Yingda Li, Zhengbo Zhang, Jing Liu, Nianfeng Liu, Yan Huang, Liang Wang arxiv

Vision-Language-Action (VLA) models leverage pretrained Vision-Language Models (VLMs) as backbones to map images and instructions to actions, demonstrating remarkable potential for generalizable robotic manipulation. To enhance performance, existing methods often incorporate extra observation cues (e.g., depth maps, point clouds) or auxiliary modules (e.g., object detectors, encoders) to enable more precise and reliable task execution, yet these typically require costly data collection and additional training. Inspired by the finding that Feed-Forward Network (FFN) in language models can act as "key-value memory", we propose Uncertainty-aware Observation Reinjection (UAOR), an effective, training-free and plug-and-play module for VLA models. Specifically, when the current language model layer exhibits high uncertainty, measured by Action Entropy, it reinjects key observation information into the next layer's Feed-Forward Network (FFN) through attention retrieval. This mechanism directly augments the hidden states with observation evidence at high-uncertainty layers, enabling more accurate and reliable action generation. Comprehensive experiments show that our method consistently improves diverse VLA models across simulation and real-world tasks with minimal overhead. Notably, UAOR eliminates the need for additional observation cues or modules, making it a versatile and practical plug-in for existing VLA pipelines. The project page is at https://uaor.jiabingyang.cn.

📄 PDF Abstract BibTeX arXiv:2602.18020

Code (0)

등록된 구현이 없습니다.

Tasks

Point Clouds

Similar Papers 제목 키워드 기반

Decision-Aware Uncertainty Evaluation of Vision-Language Model-Based Early Action Anticipation for Human-Robot Interaction

2026-03-09 · Zhaoda Du, Michael Bowman, Qiaojie Zheng, Xiaoli Zhang arxiv

Robots in shared workspaces must interpret human actions from partial, ambiguous observations, where overconfident early predictions can lead to unsafe or disruptive interaction. This challenge is amplified in egocentric…

Action AnticipationAction Recognition

Fuel savings through missed approach maneuvers based on aircraft reinjection

2022-07-07 · María Carmona, Rafael Casado, Aurelio Bermúdez, Miguel Pérez Francisco 외

Humanity is facing global challenges related to climate change, along with an energetic crisis that urgently requires optimizing any process or system able to improve global economic conditions. Taking fuel as an example…

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers

2026-02-06 · Yuxuan Yao, Yuxuan Chen, Hui Li, Kaihui Cheng 외 arxiv

Multimodal Diffusion Transformers (MMDiTs) for text-to-image generation maintain separate text and image branches, with bidirectional information flow between text tokens and visual latents throughout denoising. In this …

Text-to-Image Generation

Uncertainty-Aware Gaussian Map for Vision-Language Navigation

2026-05-26 · Jianzhe Gao, Rui Liu, Yuxuan Xu, Tongtong Cao 외 arxiv

Vision-Language Navigation (VLN) requires an agent to navigate 3D environments following natural language instructions. During navigation, existing agents commonly encounter perceptual uncertainty, such as insufficient e…

Vision-Language Navigation

MU-GeNeRF: Multi-view Uncertainty-guided Generalizable Neural Radiance Fields for Distractor-aware Scene

2026-04-20 · Wenjie Mu, Zhan Li, Chuanzhou Su, Xuanyi Shen 외 arxiv

Generalizable Neural Radiance Fields (GeNeRFs) enable high-quality scene reconstruction from sparse views and can generalize to unseen scenes. However, in real-world settings, transient distractors break cross-view struc…