paper-with-me

홈 › Papers

Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning

2026-07-06 · Matthew Foutter, Matteo Cercola, Lena Wild, Yunshan Wang, Michelle Li, Daniele Gammelli, Marco Pavone arxiv

Embodied Chain-of-Thought has emerged as a promising mechanism to enhance robot decision-making and interpretability in black-box Vision-Language Action (VLA) models. However, whether this verbalized Chain-of-Thought truthfully reflects the policy's underlying decision process remains poorly understood. We distinguish between functional reasoning, in which reasoning improves task performance, and faithful reasoning, in which reasoning truly reflects the policy's internal decision process. We argue that SoTA alignment strategies offer a necessary but insufficient notion of faithfulness, admitting reasoning whose intermediate steps can mask the causal links in action prediction through confounding factors (e.g., reasoning that is ungrounded in the environment and internally disconnected or inconsistent), restricting policy generalization. We study this gap through a human evaluation of a SoTA reasoning model for autonomous driving, revealing an inconsistent coupling between reasoning quality and downstream trajectory improvement. We then operationalize a behavioral surrogate for embodied faithfulness through a learned critic, Pinocchio, scoring observation grounding and stepwise coherence, and use this critic as a dense reward signal in post-training an embodied policy with reinforcement learning. Across withheld driving benchmarks, our post-trained planner improves faithfulness by 4% and 18% over SoTA alignment and trajectory error post-training baselines, respectively, while maintaining competitive downstream task performance. Finally, on a synthetic out-of-distribution test set, post-training for faithfulness improves policy responsiveness to rare counterfactual scenarios by 1.6x that of a SoTA policy, suggesting that faithful reasoning traces contribute to more robust, generalizable, and interpretable embodied intelligence. Project page: https://mjf-su.github.io/pinocchio/

📄 PDF Abstract BibTeX arXiv:2607.04681

Code (2)

Aaron617/agent-arXiv-daily ★ 10
BaiShuanghao/my_arXiv_daily ★ 196

Tasks

Reinforcement LearningAutonomous Driving

Similar Papers 제목 키워드 기반

Reference Games as a Testbed for the Alignment of Model Uncertainty and Clarification Requests

2026-01-12 · Manar Ali, Judith Sieker, Sina Zarrieß, Hendrik Buschmeier arxiv

In human conversation, both interlocutors play an active role in maintaining mutual understanding. When listeners are uncertain about what speakers mean, for example, they can request clarification. It is an open questio…

Teaching Vision-Language-Action Models What to See and Where to Look

2026-07-02 · Yuguang Yang, Canyu Chen, Zhewen Tan, Yizhi Wang 외 arxiv

Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing VLAs' training relies heavily on text-centric visual question answering and chain-of-thought r…

Visual Question AnsweringTrajectory PredictionAutonomous Driving

PIGLeT: Language Grounding Through Neuro-Symbolic Interaction in a 3D World

2021-06-01 · ACL 2021 5 · Rowan Zellers, Ari Holtzman, Matthew Peters, Roozbeh Mottaghi 외

We propose PIGLeT: a model that learns physical commonsense knowledge through interaction, and then uses this knowledge to ground language. We factorize PIGLeT into a physical dynamics model, and a separate language mode…

Language ModelingLanguage ModellingSentence

What Makes a Maze Look Like a Maze?

2024-09-12 · Joy Hsu, Jiayuan Mao, Joshua B. Tenenbaum, Noah D. Goodman 외

A unique aspect of human visual understanding is the ability to flexibly interpret abstract concepts: acquiring lifted rules explaining what they symbolize, grounding them across familiar and unfamiliar contexts, and mak…

Visual Reasoning

Latent Compass: Creation by Navigation

2020-12-20 · Sarah Schwettmann, Hendrik Strobelt, Mauro Martino

In Marius von Senden's Space and Sight, a newly sighted blind patient describes the experience of a corner as lemon-like, because corners "prick" sight like lemons prick the tongue. Prickliness, here, is a dimension in t…

Image Manipulation