paper-with-me

Papers

Ges3ViG : Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding

2025-01-01 · CVPR 2025 1 · Atharv Mahesh Mane, Dulanga Weerakoon, Vigneshwaran Subbaraju, Sougata Sen, Sanjay E. Sarma, Archan Misra

3-Dimensional Embodied Reference Understanding (3DERU) combines a language description and an accompanying pointing gesture to identify the most relevant target object in a 3D scene. Although prior work has explored pure language-based 3D grounding, there has been limited exploration of 3D-ERU, which also incorporates human pointing gestures. To address this gap, we introduce a data augmentation framework-Imputer, and use it to curate a new benchmark dataset-ImputeRefer for 3D-ERU, by incorporating human pointing gestures into existing 3D scene datasets that only contain language instructions. We also propose Ges3ViG, a novel model for 3D-ERU that achieves 30% improvement in accuracy as compared to other 3DERU models and 9% compared to other purely language-based 3D grounding models. Our code and dataset are available at https://github.com/AtharvMane/Ges3ViG.

📄 PDF Abstract BibTeX

Code (1)

atharvmane/ges3vig 공식 구현 pytorch

Tasks

3D visual groundingData AugmentationVisual Grounding

Similar Papers 제목 키워드 기반

Ges3ViG: Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding

2025-04-13 · Atharv Mahesh Mane, Dulanga Weerakoon, Vigneshwaran Subbaraju, Sougata Sen 외

3-Dimensional Embodied Reference Understanding (3D-ERU) combines a language description and an accompanying pointing gesture to identify the most relevant target object in a 3D scene. Although prior work has explored pur…

3D visual groundingData AugmentationVisual Grounding

Solving Visual Object Ambiguities when Pointing: An Unsupervised Learning Approach

2019-12-13 · Doreen Jirak, David Biertimpel, Matthias Kerzel, Stefan Wermter

Whenever we are addressing a specific object or refer to a certain spatial location, we are using referential or deictic gestures usually accompanied by some verbal description. Especially pointing gestures are necessary…

Objectobject-detectionObject Detection

MRPoS: Mixed Reality-Based Robot Navigation Interface Using Spatial Pointing and Speech with Large Language Model

2026-03-04 · Eduardo Iglesius, Masato Kobayashi, Yuki Uranishi arxiv

Recent advancements have made robot navigation more intuitive by transitioning from traditional 2D displays to spatially aware Mixed Reality (MR) systems. However, current MR interfaces often rely on manual "air tap" ges…

Robot Navigation

InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

2023-05-09 · Zhaoyang Liu, Yinan He, Wenhai Wang, Weiyun Wang 외

We present an interactive visual framework named InternGPT, or iGPT for short. The framework integrates chatbots that have planning and reasoning capabilities, such as ChatGPT, with non-verbal instructions like pointing …

Language Modelling

PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought

2026-06-23 · Ling Li, Bowen Liu, Zinuo Zhan, Jianhui Zhong 외 arxiv

Pointing-based visual grounding requires models to precisely locate target objects by deciphering complex spatial relationships between the visual scene and pointing gestures. Traditional methods typically encode input i…

Reinforcement LearningVisual Grounding