paper-with-me

Papers

Pointing-Based Object Recognition

2026-03-16 · Lukáš Hajdúch, Viktor Kocur arxiv

This paper presents a comprehensive pipeline for recognizing objects targeted by human pointing gestures using RGB images. As human-robot interaction moves toward more intuitive interfaces, the ability to identify targets of non-verbal communication becomes crucial. Our proposed system integrates several existing state-of-the-art methods, including object detection, body pose estimation, monocular depth estimation, and vision-language models. We evaluate the impact of 3D spatial information reconstructed from a single image and the utility of image captioning models in correcting classification errors. Experimental results on a custom dataset show that incorporating depth information significantly improves target identification, especially in complex scenes with overlapping objects. The modularity of the approach allows for deployment in environments where specialized depth sensors are unavailable.

📄 PDF Abstract BibTeX arXiv:2603.15403

Code (0)

등록된 구현이 없습니다.

Tasks

Monocular Depth EstimationObject RecognitionObject DetectionImage Captioning

Similar Papers 제목 키워드 기반

Pointing Novel Objects in Image Captioning

2019-04-25 · CVPR 2019 6 · Yehao Li, Ting Yao, Yingwei Pan, Hongyang Chao 외

Image captioning has received significant attention with remarkable improvements in recent advances. Nevertheless, images in the wild encapsulate rich knowledge and cannot be sufficiently described with models built on i…

DecoderImage CaptioningObjectObject Recognition+1

DeePoint: Visual Pointing Recognition and Direction Estimation

2023-04-14 · ICCV 2023 1 · Shu Nakamura, Yasutomo Kawanishi, Shohei Nobuhara, Ko Nishino

In this paper, we realize automatic visual recognition and direction estimation of pointing. We introduce the first neural pointing understanding method based on two key contributions. The first is the introduction of a …

What's the point? Frame-wise Pointing Gesture Recognition with Latent-Dynamic Conditional Random Fields

2015-10-20 · Christian Wittner, Boris Schauerte, Rainer Stiefelhagen

We use Latent-Dynamic Conditional Random Fields to perform skeleton-based pointing gesture classification at each time instance of a video sequence, where we achieve a frame-wise pointing accuracy of roughly 83%. Subsequ…

General ClassificationGesture Recognition

Towards Context-Aware Human-like Pointing Gestures with RL Motion Imitation

2025-09-16 · Anna Deichler, Siyang Wang, Simon Alexanderson, Jonas Beskow arxiv

Pointing is a key mode of interaction with robots, yet most prior work has focused on recognition rather than generation. We present a motion capture dataset of human pointing gestures covering diverse styles, handedness…

Reinforcement Learning

Teaching Robots Novel Objects by Pointing at Them

2020-12-25 · Sagar Gubbi Venkatesh, Raviteja Upadrashta, Shishir Kolathaya, Bharadwaj Amrutur

Robots that must operate in novel environments and collaborate with humans must be capable of acquiring new knowledge from human experts during operation. We propose teaching a robot novel objects it has not encountered …

Object