paper-with-me

Papers

Understanding Spatial Relations through Multiple Modalities

2020-07-19 · LREC 2020 5 · Soham Dan, Hangfeng He, Dan Roth

Recognizing spatial relations and reasoning about them is essential in multiple applications including navigation, direction giving and human-computer interaction in general. Spatial relations between objects can either be explicit -- expressed as spatial prepositions, or implicit -- expressed by spatial verbs such as moving, walking, shifting, etc. Both these, but implicit relations in particular, require significant common sense understanding. In this paper, we introduce the task of inferring implicit and explicit spatial relations between two entities in an image. We design a model that uses both textual and visual information to predict the spatial relations, making use of both positional and size information of objects and image embeddings. We contrast our spatial model with powerful language models and show how our modeling complements the power of these, improving prediction accuracy and coverage and facilitates dealing with unseen subjects, objects and relations.

📄 PDF Abstract BibTeX arXiv:2007.09551

Code (0)

등록된 구현이 없습니다.

Tasks

Common Sense ReasoningImplicit Relations

Similar Papers 제목 키워드 기반

Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos

2023-07-10 · CVPR 2024 1 · Sagnik Majumder, Ziad Al-Halah, Kristen Grauman

We propose a self-supervised method for learning representations based on spatial audio-visual correspondences in egocentric videos. Our method uses a masked auto-encoding framework to synthesize masked binaural (multi-c…

Active Speaker DetectionAudio DenoisingDenoising

AS3D: 2D-Assisted Cross-Modal Understanding with Semantic-Spatial Scene Graphs for 3D Visual Grounding

2025-05-07 · Feng Xiao, Hongbin Xu, Guocan Zhao, Wenxiong Kang

3D visual grounding aims to localize the unique target described by natural languages in 3D scenes. The significant gap between 3D and language modalities makes it a notable challenge to distinguish multiple similar obje…

3D visual groundingGraph AttentionObjectRelational Reasoning+1

COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence

2025-12-04 · Zefeng Zhang, Xiangzhao Hao, Hengzhu Tang, Zhenyu Zhang 외 arxiv

Visual Spatial Reasoning is crucial for enabling Multimodal Large Language Models (MLLMs) to understand object properties and spatial relationships, yet current models still struggle with 3D-aware reasoning. Existing app…

Reinforcement LearningSpatial Reasoning

Complete 3d relationships extraction modality alignment network for 3d dense captioning

2024-08-01 · IEEE Transactions on Visualization and Computer Graphics. 2024 8 · Aihua Mao, Zhi Yang, Wanxin Chen, Ran Yi 외

3D dense captioning aims to semantically describe each object detected in a 3D scene, which plays a significant role in 3D scene understanding. Previous works lack a complete definition of 3D spatial relationships and th…

3D dense captioning3D Object DetectionDense Captioningobject-detection+2

Space3D-Bench: Spatial 3D Question Answering Benchmark

2024-08-29 · Emilia Szymanska, Mihai Dusmanu, Jan-Willem Buurlage, Mahdi Rad 외

Answering questions about the spatial properties of the environment poses challenges for existing language and vision foundation models due to a lack of understanding of the 3D world notably in terms of relationships bet…

Question Answering