paper-with-me

Papers

Encoding Spatial Relations from Natural Language

2018-07-04 · Tiago Ramalho, Tomáš Kočiský, Frederic Besse, S. M. Ali Eslami, Gábor Melis, Fabio Viola, Phil Blunsom, Karl Moritz Hermann

Natural language processing has made significant inroads into learning the semantics of words through distributional approaches, however representations learnt via these methods fail to capture certain kinds of information implicit in the real world. In particular, spatial relations are encoded in a way that is inconsistent with human spatial reasoning and lacking invariance to viewpoint changes. We present a system capable of capturing the semantics of spatial relations such as behind, left of, etc from natural language. Our key contributions are a novel multi-modal objective based on generating images of scenes from their textual descriptions, and a new dataset on which to train it. We demonstrate that internal representations are robust to meaning preserving transformations of descriptions (paraphrase invariance), while viewpoint invariance is an emergent property of the system.

📄 PDF Abstract BibTeX arXiv:1807.01670

Code (1)

deepmind/slim-dataset 공식 구현 tf

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

A 2D Semantic-Aware Position Encoding for Vision Transformers

2025-05-14 · Xi Chen, Shiyang Zhou, Muqi Huang, Jiaxu Feng 외

Vision transformers have demonstrated significant advantages in computer vision tasks due to their ability to capture long-range dependencies and contextual relationships through self-attention. However, existing positio…

PositionSemantic SimilaritySemantic Textual SimilarityTranslation

Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane

2026-02-03 · Haoyu Liu, Sucheng Ren, Tingyu Zhu, Peng Wang 외 arxiv

Rotary Position Embedding (RoPE) is the de facto positional encoding in large language models due to its ability to encode relative positions and support length extrapolation. When adapted to vision transformers, the sta…

Polar Relative Positional Encoding for Video-Language Segmentation

2020-07-20 · Ke Ning, Lingxi Xie, Fei Wu, Qi Tian

In this paper, we tackle a challenging task named video-language segmentation. Given a video and a sentence in natural language, the goal is to segment the object or actor described by the sentence in video frames. To ac…

Referring Expression SegmentationSentence

Image Annotation with ISO-Space: Distinguishing Content from Structure

2014-05-01 · LREC 2014 5 · James Pustejovsky, Zachary Yocum

Natural language descriptions of visual media present interesting problems for linguistic annotation of spatial information. This paper explores the use of ISO-Space, an annotation specification to capturing spatial info…

Content-Based Image RetrievalImage CaptioningImage RetrievalSemantic Role Labeling

Watch It Twice: Video Captioning with a Refocused Video Encoder

2019-07-21 · Xiangxi Shi, Jianfei Cai, Shafiq Joty, Jiuxiang Gu

With the rapid growth of video data and the increasing demands of various applications such as intelligent video search and assistance toward visually-impaired people, video captioning task has received a lot of attentio…

Video Captioning