paper-with-me

홈 › Papers

SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation

2025-02-18 · Zekun Qi, Wenyao Zhang, Yufei Ding, Runpei Dong, Xinqiang Yu, Jingwen Li, Lingyun Xu, Baoyu Li, Xialin He, Guofan Fan, Jiazhao Zhang, JiaWei He, Jiayuan Gu, Xin Jin, Kaisheng Ma, Zhizheng Zhang, He Wang, Li Yi

Spatial intelligence is a critical component of embodied AI, promoting robots to understand and interact with their environments. While recent advances have enhanced the ability of VLMs to perceive object locations and positional relationships, they still lack the capability to precisely understand object orientations-a key requirement for tasks involving fine-grained manipulations. Addressing this limitation not only requires geometric reasoning but also an expressive and intuitive way to represent orientation. In this context, we propose that natural language offers a more flexible representation space than canonical frames, making it particularly suitable for instruction-following robotic systems. In this paper, we introduce the concept of semantic orientation, which defines object orientations using natural language in a reference-frame-free manner (e.g., the ''plug-in'' direction of a USB or the ''handle'' direction of a knife). To support this, we construct OrienText300K, a large-scale dataset of 3D models annotated with semantic orientations that link geometric understanding to functional semantics. By integrating semantic orientation into a VLM system, we enable robots to generate manipulation actions with both positional and orientational constraints. Extensive experiments in simulation and real world demonstrate that our approach significantly enhances robotic manipulation capabilities, e.g., 48.7% accuracy on Open6DOR and 74.9% accuracy on SIMPLER.

📄 PDF Abstract BibTeX arXiv:2502.13143

Code (2)

qizekun/SoFar pytorch
zhangwenyao1/open6dor_v2_execution pytorch

Tasks

Object RearrangementRobot ManipulationRobot NavigationSpatial ReasoningVisual Question Answering

Similar Papers 제목 키워드 기반

Seeing Isn't Orienting: A Cognitively Grounded Benchmark Reveals Systematic Orientation Failures in MLLMs Supplementary

2026-03-12 · Nazia Tasnim, Keanu Nichols, Yuting Yang, Nicholas Ikechukwu 외 arxiv

Humans learn object orientation progressively, from recognizing which way an object faces, to mentally rotating it, to reasoning about orientations between objects. Current vision-language benchmarks largely conflate ori…

Scene UnderstandingObject Recognition

Spatial AMR: Expanded Spatial Annotation in the Context of a Grounded Minecraft Corpus

2020-05-01 · LREC 2020 5 · Julia Bonn, Martha Palmer, Zheng Cai, Kristin Wright-Bettner

This paper presents an expansion to the Abstract Meaning Representation (AMR) annotation schema that captures fine-grained semantically and pragmatically derived spatial information in grounded corpora. We describe a new…

Abstract Meaning RepresentationMinecraftSentence

G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning

2025-11-26 · Wenbo Hu, Jingli Lin, Yilin Long, Yunlong Ran 외 arxiv

Vision-Language Models (VLMs) still lack robustness in spatial intelligence, demonstrating poor performance on spatial understanding and reasoning tasks. We attribute this gap to the absence of a visual geometry learning…

Spatial Reasoning3D Reconstruction3D scene Editing

GroundSet: A Cadastral-Grounded Dataset for Spatial Understanding with Vector Data

2026-03-15 · Roger Ferrod, Maël Lecene, Krishna Sapkota, George Leifman 외 arxiv

Precise spatial understanding in Earth Observation is essential for translating raw aerial imagery into actionable insights for critical applications like urban planning, environmental monitoring and disaster management.…

Spatial Reasoning

Cognitively-Inspired Tokens Overcome Egocentric Bias in Multimodal Models

2026-01-23 · Bridget Leonard, Scott O. Murray arxiv

Multimodal language models (MLMs) perform well on semantic vision-language tasks but fail at spatial reasoning that requires adopting another agent's visual perspective. These errors reflect a persistent egocentric bias …

Spatial Reasoning