paper-with-me

Papers

Learning to encode spatial relations from natural language

2019-05-01 · ICLR 2019 5 · Tiago Ramalho, Tomas Kocisky‎, Frederic Besse, S. M. Ali Eslami, Gabor Melis, Fabio Viola, Phil Blunsom, Karl Moritz Hermann

Natural language processing has made significant inroads into learning the semantics of words through distributional approaches, however representations learnt via these methods fail to capture certain kinds of information implicit in the real world. In particular, spatial relations are encoded in a way that is inconsistent with human spatial reasoning and lacking invariance to viewpoint changes. We present a system capable of capturing the semantics of spatial relations such as behind, left of, etc from natural language. Our key contributions are a novel multi-modal objective based on generating images of scenes from their textual descriptions, and a new dataset on which to train it. We demonstrate that internal representations are robust to meaning preserving transformations of descriptions (paraphrase invariance), while viewpoint invariance is an emergent property of the system.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Representing Spatial Relations in FrameNet

2018-06-01 · WS 2018 6 · Miriam R. L. Petruck, Michael J. Ellsworth

While humans use natural language to express spatial relations between and across entities in the world with great facility, natural language systems have a facility that depends on that human facility. This position pap…

Position

Encoding Spatial Relations from Natural Language

2018-07-04 · Tiago Ramalho, Tomáš Kočiský, Frederic Besse, S. M. Ali Eslami 외

Natural language processing has made significant inroads into learning the semantics of words through distributional approaches, however representations learnt via these methods fail to capture certain kinds of informati…

Spatial Reasoning

The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models

2026-03-23 · Kelly Cui, Nikhil Prakash, Shoval Messica, Ayush Raina 외 arxiv

Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their properties and spatial relations. Yet it remains unclear where and how such …

Visual Question AnsweringImage Captioning

GeoLM: Empowering Language Models for Geospatially Grounded Language Understanding

2023-10-23 · Zekun Li, Wenxuan Zhou, Yao-Yi Chiang, Muhao Chen

Humans subconsciously engage in geospatial reasoning when reading articles. We recognize place names and their spatial relations in text and mentally associate them with their physical locations on Earth. Although pretra…

ArticlesContrastive LearningEntity TypingLanguage Modeling+4

SpatiaLoc: Leveraging Multi-Level Spatial Enhanced Descriptors for Cross-Modal Localization

2026-01-07 · Tianyi Shang, Pengjie Xu, Zhaojun Deng, Zhenyu Li 외 arxiv

Cross-modal localization using text and point clouds enables robots to localize themselves via natural language descriptions, with applications in autonomous navigation and interaction between humans and robots. In this …

Point Clouds