paper-with-me

Papers

Embodied4C: Measuring What Matters for Embodied Vision-Language Navigation

2025-12-19 · Tin Stribor Sohn, Maximilian Dillitzer, Jason J. Corso, Eric Sax arxiv

Vision-language navigation requires agents to reason and act under constraints of embodiment. While vision-language models (VLMs) demonstrate strong generalization, current benchmarks provide limited understanding of how embodiment -- i.e., the choice of physical platform, sensor configuration, and modality alignment -- influences perception, reasoning, and control. We introduce Embodied4C, a closed-loop benchmark designed as a Turing test for embodied reasoning. The benchmark evaluates the core embodied capabilities of VLMs across three heterogeneous embodiments -- autonomous vehicles, aerial drones, and robotic manipulators -- through approximately 1.1K one-shot reasoning questions and 58 goal-directed navigation tasks. These tasks jointly assess four foundational dimensions: semantic, spatial, temporal, and physical reasoning. Each embodiment presents dynamic sensor configurations and environment variations to probe generalization beyond platform-specific adaptation. To prevent embodiment overfitting, Embodied4C integrates domain-far queries targeting abstract and cross-context reasoning. Comprehensive evaluation across ten state-of-the-art VLMs and four embodied control baselines shows that cross-modal alignment and instruction tuning matter more than scale, while spatial and temporal reasoning remains the primary bottleneck for reliable embodied competence.

📄 PDF Abstract BibTeX arXiv:2512.18028

Code (0)

등록된 구현이 없습니다.

Tasks

Vision-Language NavigationAutonomous Vehicles

Similar Papers 제목 키워드 기반

Guava: An Effective and Universal Harness for Embodied Manipulation

2026-06-16 · Haowen Liu, Xirui Li, Shaoxiong Yao, Peng Shi 외 arxiv

Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. Harnessing models through embodied tools use offers a promising alternative to end-to-end vision-languag…

Embodied Question Answering

2017-11-30 · CVPR 2018 6 · Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee 외

We present a new AI task -- Embodied Question Answering (EmbodiedQA) -- where an agent is spawned at a random location in a 3D environment and asked a question ("What color is the car?"). In order to answer, the agent mu…

Embodied Question AnsweringNavigateQuestion Answeringreinforcement-learning+2

ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion

2023-02-09 · Ying Shen, Daniel Bis, Cynthia Lu, Ismini Lourentzou

The research community has shown increasing interest in designing intelligent embodied agents that can assist humans in accomplishing tasks. Although there have been significant advancements in related vision-language be…

Question AnsweringVisual Navigation

ERQA-Plus: A Diagnostic Benchmark for Reasoning in Embodied AI

2026-06-16 · Hong Yang, Basura Fernando arxiv

Generalist embodied agents require more than object recognition: they must reason about spatial relations, actions, procedures, human intentions, environmental constraints, and commonsense consequences from situated visu…

Question GenerationObject RecognitionQuestion AnsweringSpatial Reasoning

Improving Vision-and-Language Navigation with Image-Text Pairs from the Web

2020-04-30 · ECCV 2020 8 · Arjun Majumdar, Ayush Shrivastava, Stefan Lee, Peter Anderson 외

Following a navigation instruction such as 'Walk down the stairs and stop at the brown sofa' requires embodied AI agents to ground scene elements referenced via language (e.g. 'stairs') to visual content in the environme…

Vision and Language Navigation