paper-with-me

홈 › Papers

M3T: Discrete Multi-Modal Motion Tokens for Sign Language Production

2026-03-24 · Alexandre Symeonidis-Herzig, Jianhe Low, Ozge Mercanoglu Sincan, Richard Bowden arxiv

Sign language production requires more than hand motion generation. Non-manual features, including mouthings, eyebrow raises, gaze, and head movements, are grammatically obligatory and cannot be recovered from manual articulators alone. Existing 3D production systems face two barriers to integrating them: the standard body model provides a facial space too low-dimensional to encode these articulations, and when richer representations are adopted, standard discrete tokenization suffers from codebook collapse, leaving most of the expression space unreachable. We propose SMPL-FX, which couples FLAME's rich expression space with the SMPL-X body, and tokenize the resulting representation with modality-specific Finite Scalar Quantization VAEs for body, hands, and face. M3T is an autoregressive transformer trained on this multi-modal motion vocabulary, with an auxiliary translation objective that encourages semantically grounded embeddings. Across three standard benchmarks (How2Sign, CSL-Daily, Phoenix14T) M3T achieves state-of-the-art sign language production quality, and on NMFs-CSL, where signs are distinguishable only by non-manual features, reaches 58.3% accuracy against 49.0% for the strongest comparable pose baseline.

📄 PDF Abstract BibTeX arXiv:2603.23617

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Unified Framework for Multimodal, Multi-Part Human Motion Synthesis

2023-11-28 · Zixiang Zhou, Yu Wan, Baoyuan Wang

The field has made significant progress in synthesizing realistic human motion driven by various modalities. Yet, the need for different methods to animate various body parts according to different control signals limits…

Motion GenerationMotion Synthesis

MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators

2023-06-19 · Yaqi Zhang, Di Huang, Bin Liu, Shixiang Tang 외

Generating realistic human motion from given action descriptions has experienced significant advancements because of the emerging requirement of digital humans. While recent works have achieved impressive results in gene…

Motion Generation

HSI-GPT: A General-Purpose Large Scene-Motion-Language Model for Human Scene Interaction

2025-01-01 · CVPR 2025 1 · YuAn Wang, YaLi Li, Xiang Li, Shengjin Wang

While flourishing developments have been witnessed in text-to-motion generation, synthesizing physically realistic, controllable, language-conditioned Human Scene Interactions (HSI) remains a relatively underexplored…

DescriptiveInstruction FollowingLanguage ModelingLanguage Modelling+1

AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis

2026-07-17 · Zhenqi Jia, Yuan Zhao, Aruukhan, Rui Liu 외 arxiv

Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods struggle to render authentic human emotions…

Speech Synthesis

TM2T: Stochastic and Tokenized Modeling for the Reciprocal Generation of 3D Human Motions and Texts

2022-07-04 · Chuan Guo, Xinxin Zuo, Sen Wang, Li Cheng

Inspired by the strong ties between vision and language, the two intimate human sensing and communication modalities, our paper aims to explore the generation of 3D human full-body motions from texts, as well as its reci…

Machine TranslationMotion CaptioningMotion SynthesisNMT