paper-with-me

Papers

FloorplanQA: A Benchmark for Spatial Reasoning in LLMs using Structured Representations

2025-07-10 · Fedor Rodionov, Abdelrahman Eldesokey, Michael Birsak, John Femiani, Bernard Ghanem, Peter Wonka arxiv

We introduce FloorplanQA, a diagnostic benchmark for evaluating spatial reasoning in large language models (LLMs). FloorplanQA is grounded in structured representations of indoor scenes, such as (e.g., kitchens, living rooms, bedrooms, bathrooms, and others), encoded symbolically in JSON or XML layouts. The benchmark covers core spatial tasks, including distance measurement, visibility, path finding, and object placement within constrained spaces. Our results across a variety of frontier open-source and commercial LLMs reveal that while models may succeed in shallow queries, they often fail to respect physical constraints, preserve spatial coherence, though they remain mostly robust to small spatial perturbations. FloorplanQA uncovers a blind spot in today's LLMs: inconsistent reasoning about indoor layouts. We hope this benchmark inspires new work on language models that can accurately infer and manipulate spatial and geometric properties in practical settings.

📄 PDF Abstract BibTeX arXiv:2507.07644

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Think Visually: Question Answering through Virtual Imagery

2018-05-25 · ACL 2018 7 · Ankit Goyal, Jian Wang, Jia Deng

In this paper, we study the problem of geometric reasoning in the context of question-answering. We introduce Dynamic Spatial Memory Network (DSMN), a new deep network architecture designed for answering questions that a…

Question AnsweringVisual Commonsense Reasoning

Spatial4D-Bench: A Versatile 4D Spatial Intelligence Benchmark

2025-12-31 · Pan Wang, Yang Liu, Guile Wu, Eduardo R. Corral-Soto 외 arxiv

4D spatial intelligence involves perceiving and processing how objects move or change over time. Humans naturally possess 4D spatial intelligence, supporting a broad spectrum of spatial reasoning abilities. To what exten…

Scene UnderstandingAction RecognitionSpatial Reasoning

3D-Layout-R1: Structured Reasoning for Language-Instructed Spatial Editing

2026-03-23 · Haoyu Zhen, Xiaolong Li, Yilin Zhao, Han Zhang 외 arxiv

Large Language Models (LLMs) and Vision Language Models (VLMs) have shown impressive reasoning abilities, yet they struggle with spatial understanding and layout consistency when performing fine-grained visual editing. W…

SSR3D-LLM: Structured Spatial Reasoning via Latent Steps for Fine-Grained Grounding in Unified 3D-LLMs

2026-05-27 · Jiawei Li, Ziyi Liu, Weijie Shi, Long Chen 외 arxiv

3D object grounding localizes referred objects in a 3D scene from natural language. Unified instance-centric 3D-LLMs aim to solve grounding together with dialog, QA, and captioning, yet many rely on a single pointer-styl…

Spatial Reasoning

EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks

2025-03-14 · Yi Zhang, Qiang Zhang, Xiaozhu Ju, Zhaoyang Liu 외

While multimodal large language models (MLLMs) have made groundbreaking progress in embodied intelligence, they still face significant challenges in spatial reasoning for complex long-horizon tasks. To address this gap, …

Spatial Reasoning