paper-with-me

Papers

Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models

2026-07-08 · Kaito Watanabe, Taisei Yamamoto, Tomoki Doi, Hitomi Yanaka arxiv

One of the expected abilities of vision-language models (VLMs) is spatial reasoning ability based on a given text and image. To evaluate the spatial reasoning abilities of VLMs, we focus on the use of spatial deictic expressions, which are defined as spatial expressions whose referent is determined by their situational context, such as `this'' and `that''. To handle spatial deictic expressions, VLMs must jointly reason over language and visual space, grounding context-dependent references in the image's spatial structure. In addition, selecting appropriate spatial deictic expressions across languages requires VLMs to understand the language-specific spatial distinctions encoded by these expressions. In this paper, we develop a benchmark to evaluate the multilingual ability of VLMs to use spatial deictic expressions in four languages. Our experiments using this benchmark reveal that the tested models use demonstratives in a manner different from that of humans, particularly in selecting the appropriate demonstratives based on the distance to the object.

📄 PDF Abstract BibTeX arXiv:2607.07251

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

DeicticVLA: Unifying Instruction Modes Based on Language and Deictic Gestures in a Single VLA

2026-08-28 · Kango Yanagida, Tatsuya Aoki, Yuichiro Yoshikawa, Takato Horii arxiv

Vision-Language-Action models (VLAs) allow users to specify manipulation tasks in natural language, but distinguishing a target or placement goal among objects of the same category or similar appearance requires detailed…

Investigating Information-Theoretic Properties of the Typology of Spatial Demonstratives

2022-07-01 · NAACL (SIGTYP) 2022 7 · Sihan Chen, Richard Futrell, Kyle Mahowald

Using data from Nintemann et al. (2020), we explore the variability in complexity and informativity across spatial demonstrative systems using spatial deictic lexicons from 223 languages. We argue from an information-the…

DeiSAM: Segment Anything with Deictic Prompting

2024-02-21 · Hikaru Shindo, Manuel Brack, Gopika Sudhakaran, Devendra Singh Dhami 외

Large-scale, pre-trained neural networks have demonstrated strong capabilities in various tasks, including zero-shot image segmentation. To identify concrete objects in complex scenes, humans instinctively rely on deicti…

Image SegmentationSegmentationSemantic Segmentation

Comprehension of Multilingual Expressions Referring to Target Objects in Visual Inputs

2025-11-14 · Francisco Nogueira, Alexandre Bernardino, Bruno Martins arxiv

Referring Expression Comprehension (REC) requires models to localize objects in images based on natural language descriptions. Research on the area remains predominantly English-centric, despite increasing global deploym…

Referring ExpressionMachine TranslationVisual Grounding

Listening with the Eyes: Benchmarking Egocentric Co-Speech Grounding across Space and Time

2026-03-09 · Weijie Zhou, Xuantang Xiong, Zhenlin Hu, Xiaomeng Zhu 외 arxiv

In situated collaboration, speakers often use intentionally underspecified deictic commands (e.g., ``pass me \textit{that}''), whose referent becomes identifiable only by aligning speech with a brief co-speech pointing \…