paper-with-me

홈 › Papers

SPARE3D: A Dataset for SPAtial REasoning on Three-View Line Drawings

2020-03-31 · CVPR 2020 6 · Wenyu Han, Siyuan Xiang, Chenhui Liu, Ruoyu Wang, Chen Feng

Spatial reasoning is an important component of human intelligence. We can imagine the shapes of 3D objects and reason about their spatial relations by merely looking at their three-view line drawings in 2D, with different levels of competence. Can deep networks be trained to perform spatial reasoning tasks? How can we measure their "spatial intelligence"? To answer these questions, we present the SPARE3D dataset. Based on cognitive science and psychometrics, SPARE3D contains three types of 2D-3D reasoning tasks on view consistency, camera pose, and shape generation, with increasing difficulty. We then design a method to automatically generate a large number of challenging questions with ground truth answers for each task. They are used to provide supervision for training our baseline models using state-of-the-art architectures like ResNet. Our experiments show that although convolutional networks have achieved superhuman performance in many visual learning tasks, their spatial reasoning performance on SPARE3D tasks is either lower than average human performance or even close to random guesses. We hope SPARE3D can stimulate new problem formulations and network designs for spatial reasoning to empower intelligent robots to operate effectively in the 3D world via 2D sensors. The dataset and code are available at https://ai4ce.github.io/SPARE3D.

📄 PDF Abstract BibTeX arXiv:2003.14034

Code (1)

ai4ce/SPARE3D 공식 구현 pytorch

Tasks

Spatial Reasoning

Methods 이 논문이 사용한 방법론

Average Pooling 설명 없음
Global Average Pooling Global Average Pooling is a pooling operation designed to replace fully connected layers in classical CNNs. The idea is to generate one feature map for each corresponding…
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Batch Normalization 설명 없음
Bottleneck Residual Block A Bottleneck Residual Block is a variant of the residual block that utilises 1x1 convolutions to create a bottleneck. The…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Kaiming Initialization 설명 없음

Similar Papers 제목 키워드 기반

Self-supervised Spatial Reasoning on Multi-View Line Drawings

2021-04-27 · CVPR 2022 1 · Siyuan Xiang, Anbang Yang, Yanfei Xue, Yaoqing Yang 외

Spatial reasoning on multi-view line drawings by state-of-the-art supervised deep networks is recently shown with puzzling low performances on the SPARE3D dataset. Based on the fact that self-supervised learning is helpf…

Binary ClassificationContrastive LearningMulti-class ClassificationSelf-Supervised Learning+1

Learning Multi-View Spatial Reasoning from Cross-View Relations

2026-03-30 · Suchae Jeong, Jaehwi Song, Haeone Lee, Hanna Kim 외 arxiv

Vision-language models (VLMs) have achieved impressive results on single-view vision tasks, but lack the multi-view spatial reasoning capabilities essential for embodied AI systems to understand 3D environments and manip…

Spatial Reasoning

CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark

2026-05-18 · Wei Wang, Yuqian Yuan, Tianwei Lin, Wenqiao Zhang 외 arxiv

Spatial intelligence requires multimodal large language models (MLLMs) to move beyond single-view perception and reason consistently about objects, visibility, geometry, and interactions across multiple viewpoints. Howev…

Spatial Reasoning

TransBiolab: A Real-World Multi-View Dataset of Cluttered Transparent Biomedical Objects

2026-07-23 · Ke Ma, Yifei Wang, Meng Wang, Tian Xia arxiv

Autonomous biomedical laboratories increasingly rely on visual perception to recognize, localize, and manipulate transparent plasticware, yet high-quality real-world datasets for this setting remain limited. The scarcity…

Robot Manipulation6D Pose EstimationDepth Estimation

How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective

2025-09-23 · Songsong Yu, Yuxin Chen, Hao Ju, Lianjie Jia 외 arxiv

Visual Spatial Reasoning (VSR) is a core human cognitive ability and a critical requirement for advancing embodied intelligence and autonomous systems. Despite recent progress in Vision-Language Models (VLMs), achieving …

Spatial Reasoning