paper-with-me

홈 › Papers

Spatial AMR: Expanded Spatial Annotation in the Context of a Grounded Minecraft Corpus

2020-05-01 · LREC 2020 5 · Julia Bonn, Martha Palmer, Zheng Cai, Kristin Wright-Bettner

This paper presents an expansion to the Abstract Meaning Representation (AMR) annotation schema that captures fine-grained semantically and pragmatically derived spatial information in grounded corpora. We describe a new lexical category conceptualization and set of spatial annotation tools built in the context of a multimodal corpus consisting of 170 3D structure-building dialogues between a human architect and human builder in Minecraft. Minecraft provides a particularly beneficial spatial relation-elicitation environment because it automatically tracks locations and orientations of objects and avatars in the space according to an absolute Cartesian coordinate system. Through a two-step process of sentence-level and document-level annotation designed to capture implicit information, we leverage these coordinates and bearings in the AMRs in combination with spatial framework annotation to ground the spatial language in the dialogues to absolute space.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Abstract Meaning RepresentationMinecraftSentence

Similar Papers 제목 키워드 기반

G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning

2025-11-26 · Wenbo Hu, Jingli Lin, Yilin Long, Yunlong Ran 외 arxiv

Vision-Language Models (VLMs) still lack robustness in spatial intelligence, demonstrating poor performance on spatial understanding and reasoning tasks. We attribute this gap to the absence of a visual geometry learning…

Spatial Reasoning3D Reconstruction3D scene Editing

Spatial-Conditioned Reasoning in Long-Egocentric Videos

2026-01-26 · James Tribble, Hao Wang, Si-En Hong, Chaoyi Zhou 외 arxiv

Long-horizon egocentric video presents significant challenges for visual navigation due to viewpoint drift and the absence of persistent geometric context. Although recent vision-language models perform well on image and…

Spatial ReasoningVisual Navigation

AutoTour: Automatic Photo Tour Guide with Smartphones and LLMs

2026-01-11 · Huatao Xu, Zihe Liu, Zilin Zeng, Baichuan Li 외 arxiv

We present AutoTour, a system that enhances user exploration by automatically generating fine-grained landmark annotations and descriptive narratives for photos captured by users. The key idea of AutoTour is to fuse visu…

Geometric Matching

Beyond a Single Frame: Multi-Frame Spatially Grounded Reasoning Across Volumetric MRI

2026-04-17 · Lama Moukheiber, Caleb M. Yeung, Haotian Xue, Alec Helbling 외 arxiv

Spatial reasoning and visual grounding are core capabilities for vision-language models (VLMs), yet most medical VLMs produce predictions without transparent reasoning or spatial evidence. Existing benchmarks also evalua…

Visual Question AnsweringSpatial ReasoningVisual Grounding

UrbanGraphEmbeddings: Learning and Evaluating Spatially Grounded Multimodal Embeddings for Urban Science

2026-02-09 · Jie Zhang, Xingtong Yu, Yuan Fang, Rudi Stouffs 외 arxiv

Learning transferable multimodal embeddings for urban environments is challenging because urban understanding is inherently spatial, yet existing datasets and benchmarks lack explicit alignment between street-view images…

Contrastive LearningSpatial ReasoningImage Retrieval