paper-with-me

홈 › Papers

TransforMerger: Transformer-based Voice-Gesture Fusion for Robust Human-Robot Communication

2025-04-02 · Petr Vanc, Karla Stepanova

As human-robot collaboration advances, natural and flexible communication methods are essential for effective robot control. Traditional methods relying on a single modality or rigid rules struggle with noisy or misaligned data as well as with object descriptions that do not perfectly fit the predefined object names (e.g. 'Pick that red object'). We introduce TransforMerger, a transformer-based reasoning model that infers a structured action command for robotic manipulation based on fused voice and gesture inputs. Our approach merges multimodal data into a single unified sentence, which is then processed by the language model. We employ probabilistic embeddings to handle uncertainty and we integrate contextual scene understanding to resolve ambiguous references (e.g., gestures pointing to multiple objects or vague verbal cues like "this"). We evaluate TransforMerger in simulated and real-world experiments, demonstrating its robustness to noise, misalignment, and missing information. Our results show that TransforMerger outperforms deterministic baselines, especially in scenarios requiring more contextual knowledge, enabling more robust and flexible human-robot communication. Code and datasets are available at: http://imitrob.ciirc.cvut.cz/publications/transformerger.

📄 PDF Abstract BibTeX arXiv:2504.01708

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingObjectScene UnderstandingSentence

Similar Papers 제목 키워드 기반

DeepGesture: A conversational gesture synthesis system based on emotions and semantics

2025-07-03 · Thanh Hoang-Minh

Along with the explosion of large language models, improvements in speech synthesis, advancements in hardware, and the evolution of computer graphics, the current bottleneck in creating digital humans lies in generating …

Gesture GenerationMotion SynthesisSpeech SynthesisUnity

CoCoGesture: Toward Coherent Co-speech 3D Gesture Generation in the Wild

2024-05-27 · Xingqun Qi, Hengyuan Zhang, Yatian Wang, Jiahao Pan 외

Deriving co-speech 3D gestures has seen tremendous progress in virtual avatar animation. Yet, the existing methods often produce stiff and unreasonable gestures with unseen human speech inputs due to the limited 3D speec…

Gesture Generation

Efficient Multimodal Neural Networks for Trigger-less Voice Assistants

2023-05-20 · Sai Srujana Buddi, Utkarsh Oggy Sarawgi, Tashweena Heeramun, Karan Sawnhey 외

The adoption of multimodal interactions by Voice Assistants (VAs) is growing rapidly to enhance human-computer interactions. Smartwatches have now incorporated trigger-less methods of invoking VAs, such as Raise To Speak…

Decision Making

Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers

2025-03-13 · Yasheng Sun, Zhiliang Xu, Hang Zhou, Jiazhi Guan 외

Co-speech gesture video synthesis is a challenging task that requires both probabilistic modeling of human gestures and the synthesis of realistic images that align with the rhythmic nuances of speech. To address these c…

Taming Diffusion Models for Audio-Driven Co-Speech Gesture Generation

2023-03-16 · CVPR 2023 1 · Lingting Zhu, Xian Liu, Xuanyu Liu, Rui Qian 외

Animating virtual avatars to make co-speech gestures facilitates various applications in human-machine interaction. The existing methods mainly rely on generative adversarial networks (GANs), which typically suffer from …

DiversityGesture Generation