paper-with-me

Papers

DeicticVLA: Unifying Instruction Modes Based on Language and Deictic Gestures in a Single VLA

2026-08-28 · Kango Yanagida, Tatsuya Aoki, Yuichiro Yoshikawa, Takato Horii arxiv

Vision-Language-Action models (VLAs) allow users to specify manipulation tasks in natural language, but distinguishing a target or placement goal among objects of the same category or similar appearance requires detailed expressions that VLAs may not use reliably. We propose DeicticVLA, which canonicalizes Language Instruction (LI), Vision-Language Instruction (VLI), and Visual Instruction (VI) into a text prompt and deictic masks through text-prompt completion and deictic gesture grounding, enabling a single pretrained VLA to handle all three instruction modes. With a shared backbone, demonstrations, and matched training steps, we compare two RGB visual prompting methods, two separate-channel mask prompting methods, and three training strategies in simulation. Under two-stage training, the four prompting methods achieve high in-distribution success but differ in their ability to use deictic masks in unseen layouts. Across methods, training-strategy ablations show that two-stage training improves such use, while retaining second-stage LI data mitigates forgetting without reducing VLI and VI performance. In three real-world tasks, one policy supports all modes. VLI and VI outperform LI under unseen expressions, appearance changes, and novel objects. For unseen categories, both achieve 100% success, compared with 16.7% for jointly trained LI. These results demonstrate the unified three-mode interface and guide DeicticVLA design.

📄 PDF Abstract BibTeX arXiv:2608.28108

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DeiSAM: Segment Anything with Deictic Prompting

2024-02-21 · Hikaru Shindo, Manuel Brack, Gopika Sudhakaran, Devendra Singh Dhami 외

Large-scale, pre-trained neural networks have demonstrated strong capabilities in various tasks, including zero-shot image segmentation. To identify concrete objects in complex scenes, humans instinctively rely on deicti…

Image SegmentationSegmentationSemantic Segmentation

Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models

2026-07-08 · Kaito Watanabe, Taisei Yamamoto, Tomoki Doi, Hitomi Yanaka arxiv

One of the expected abilities of vision-language models (VLMs) is spatial reasoning ability based on a given text and image. To evaluate the spatial reasoning abilities of VLMs, we focus on the use of spatial deictic exp…

Spatial Reasoning

Investigating Information-Theoretic Properties of the Typology of Spatial Demonstratives

2022-07-01 · NAACL (SIGTYP) 2022 7 · Sihan Chen, Richard Futrell, Kyle Mahowald

Using data from Nintemann et al. (2020), we explore the variability in complexity and informativity across spatial demonstrative systems using spatial deictic lexicons from 223 languages. We argue from an information-the…

Age-Related Differences in the Perception of Eye-Gaze from a Social Robot

2026-03-09 · Lucas Morillo-Mendez, Martien G. S. Schrooten, Oscar Martinez Mozos arxiv

There is an increasing interest in social robots assisting older adults during daily life tasks. In this context, non-verbal cues such as deictic gaze are important in natural communication in human-robot interaction. Ho…

How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing

2026-02-02 · Huanyu Zhang, Xuehai Bai, Chengzu Li, Chen Liang 외 arxiv

Recent generative models have achieved remarkable progress in image editing. However, existing systems and benchmarks remain largely text-guided. In contrast, human communication is inherently multimodal, where visual in…

visual instruction followingImage Editing