paper-with-me

홈 › Papers

Towards Embodied Cognition in Robots via Spatially Grounded Synthetic Worlds

2025-05-20 · Joel Currie, Gioele Migno, Enrico Piacenti, Maria Elena Giannaccini, Patric Bach, Davide De Tommaso, Agnieszka Wykowska

We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability for embodied cognition essential for Human-Robot Interaction (HRI). As a first step toward this goal, we introduce a synthetic dataset, generated in NVIDIA Omniverse, that enables supervised learning for spatial reasoning tasks. Each instance includes an RGB image, a natural language description, and a ground-truth 4X4 transformation matrix representing object pose. We focus on inferring Z-axis distance as a foundational skill, with future extensions targeting full 6 Degrees Of Freedom (DOFs) reasoning. The dataset is publicly available to support further research. This work serves as a foundational step toward embodied AI systems capable of spatial understanding in interactive human-robot scenarios.

📄 PDF Abstract BibTeX arXiv:2505.14366

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Grounded Gesture Generation: Language, Motion, and Space

2025-07-06 · Anna Deichler, Jim O'Regan, Teo Guichoux, David Johansson 외 arxiv

Human motion generation has advanced rapidly in recent years, yet the critical problem of creating spatially grounded, context-aware gestures has been largely overlooked. Existing models typically specialize either in de…

Synthetic Data GenerationGesture Generation

Evaluation of Vision-LLMs in Surveillance Video

2025-10-27 · Pascal Benschop, Cristian Meo, Justin Dauwels, Jelte P. Mense arxiv

The widespread use of cameras in our society has created an overwhelming amount of video data, far exceeding the capacity for human monitoring. This presents a critical challenge for public safety and security, as the ti…

Action RecognitionSpatial Reasoning

Beyond Episodic Evaluation: Memory Architectural Bottlenecks in Sequential Embodied Question Answering

2026-07-23 · Zikui Cai, Kaushal Janga, Tan Dat Dao, Seungjae Lee 외 arxiv

Embodied question answering (EQA) is traditionally evaluated under an episodic formulation, where agents solve each task independently and reset internal state between episodes. However, real-world robots operate continu…

Question Answering

Grounded Decoding: Guiding Text Generation with Grounded Models for Embodied Agents

2023-03-01 · NeurIPS 2023 11 · Wenlong Huang, Fei Xia, Dhruv Shah, Danny Driess 외

Recent progress in large language models (LLMs) has demonstrated the ability to learn and leverage Internet-scale knowledge through pre-training with autoregressive models. Unfortunately, applying such models to settings…

Language ModelingLanguage ModellingText Generation

Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation

2024-09-09 · Muraleekrishna Gopinathan, Martin Masek, Jumana Abu-Khalaf, David Suter

Embodied AI aims to develop robots that can \textit{understand} and execute human language instructions, as well as communicate in natural languages. On this front, we study the task of generating highly detailed navigat…

Vision and Language Navigation