Prompt-to-Gesture: Measuring the Capabilities of Image-to-Video Deictic Gesture Generation
Gesture recognition research, unlike NLP, continues to face acute data scarcity, with progress constrained by the need for costly human recordings or image processing approaches that cannot generate authentic variability in the gestures themselves. Recent advancements in image-to-video foundation models have enabled the generation of photorealistic, semantically rich videos guided by natural language. These capabilities open up new possibilities for creating effort-free synthetic data, raising the critical question of whether video Generative AI models can augment and complement traditional human-generated gesture data. In this paper, we introduce and analyze prompt-based video generation to construct a realistic deictic gestures dataset and rigorously evaluate its effectiveness for downstream tasks. We propose a data generation pipeline that produces deictic gestures from a small number of reference samples collected from human participants, providing an accessible approach that can be leveraged both within and beyond the machine learning community. Our results demonstrate that the synthetic gestures not only align closely with real ones in terms of visual fidelity but also introduce meaningful variability and novelty that enrich the original data, further supported by superior performance of various deep models using a mixed dataset. These findings highlight that image-to-video techniques, even in their early stages, offer a powerful zero-shot approach to gesture synthesis with clear benefits for downstream tasks.
Code (0)
등록된 구현이 없습니다.
Tasks
Gesture RecognitionGesture GenerationVideo GenerationSimilar Papers 제목 키워드 기반
Zero-shot Prompt-based Video Encoder for Surgical Gesture Recognition
Purpose: In order to produce a surgical gesture recognition system that can support a wide variety of procedures, either a very large annotated dataset must be acquired, or fitted models must generalize to new labels (so…
Gesture RecognitionSurgical Gesture RecognitionFoodTrack: Estimating Handheld Food Portions with Egocentric Video
Accurately tracking food consumption is crucial for nutrition and health monitoring. Traditional approaches typically require specific camera angles, non-occluded images, or rely on gesture recognition to estimate intake…
Gesture RecognitionNutritionMotion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models
The automatic generation of controllable co-speech gestures has recently gained growing attention. While existing systems typically achieve gesture control through predefined categorical labels or implicit pseudo-labels …
Gesture GenerationThink Step by Step: Chain-of-Gesture Prompting for Error Detection in Robotic Surgical Videos
Despite significant advancements in robotic systems and surgical data science, ensuring safe and optimal execution in robot-assisted minimally invasive surgery (RMIS) remains a complex challenge. Current surgical error d…
Temporal Information ExtractionVisual ReasoningAudio-Driven Co-Speech Gesture Video Generation
Co-speech gesture is crucial for human-machine interaction and digital entertainment. While previous works mostly map speech audio to human skeletons (e.g., 2D keypoints), directly generating speakers' gestures in the im…
Video Generation