paper-with-me

홈 › Papers

Pointing-Guided Target Estimation via Transformer-Based Attention

2025-09-05 · Luca Müller, Hassan Ali, Philipp Allgeuer, Lukáš Gajdošech, Stefan Wermter arxiv

Deictic gestures, like pointing, are a fundamental form of non-verbal communication, enabling humans to direct attention to specific objects or locations. This capability is essential in Human-Robot Interaction (HRI), where robots should be able to predict human intent and anticipate appropriate responses. In this work, we propose the Multi-Modality Inter-TransFormer (MM-ITF), a modular architecture to predict objects in a controlled tabletop scenario with the NICOL robot, where humans indicate targets through natural pointing gestures. Leveraging inter-modality attention, MM-ITF maps 2D pointing gestures to object locations, assigns a likelihood score to each, and identifies the most likely target. Our results demonstrate that the method can accurately predict the intended object using monocular RGB data, thus enabling intuitive and accessible human-robot collaboration. To evaluate the performance, we introduce a patch confusion matrix, providing insights into the model's predictions across candidate object locations. Code available at: https://github.com/lucamuellercode/MMITF.

📄 PDF Abstract BibTeX arXiv:2509.05031

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DeePoint: Visual Pointing Recognition and Direction Estimation

2023-04-14 · ICCV 2023 1 · Shu Nakamura, Yasutomo Kawanishi, Shohei Nobuhara, Ko Nishino

In this paper, we realize automatic visual recognition and direction estimation of pointing. We introduce the first neural pointing understanding method based on two key contributions. The first is the introduction of a …

Networked pointing system: Bearing-only target localization and pointing control

2025-06-23 · Shiyao Li, Bo Zhu, Yining Zhou, Jie Ma 외

In the paper, we formulate the target-pointing consensus problem where the headings of agents are required to point at a common target. Only a few agents in the network can measure the bearing information of the target. …

VistaRef: Boosting Visual Spatial Orientation Awareness for Pointing-to-Object Detection

2026-06-23 · Ling Li, Zhizhen Cai, Xinkun Wu, Ziyu Zhu 외 arxiv

Grounding deictic gestures in natural images is fundamental to AR and human-robot collaboration, providing a basis for seamless spatial interaction. While Transformer-based visual models have achieved significant progres…

Object Detection

The Anatomy of an Edit: Mechanism-Guided Activation Steering for Knowledge Editing

2026-03-21 · Yuan Cao, Mingyang Wang, Hinrich Schütze arxiv

Large language models (LLMs) are increasingly used as knowledge bases, but keeping them up to date requires targeted knowledge editing (KE). However, it remains unclear how edits are implemented inside the model once app…

knowledge editing

Pose-Oriented Transformer with Uncertainty-Guided Refinement for 2D-to-3D Human Pose Estimation

2023-02-15 · Han Li, Bowen Shi, Wenrui Dai, Hongwei Zheng 외

There has been a recent surge of interest in introducing transformers to 3D human pose estimation (HPE) due to their powerful capabilities in modeling long-term dependencies. However, existing transformer-based methods t…

3D Human Pose EstimationPose EstimationPosition