MotionScript: Natural Language Descriptions for Expressive 3D Human Motions
We introduce MotionScript, a novel framework for generating highly detailed, natural language descriptions of 3D human motions. Unlike existing motion datasets that rely on broad action labels or generic captions, MotionScript provides fine-grained, structured descriptions that capture the full complexity of human movement including expressive actions (e.g., emotions, stylistic walking) and interactions beyond standard motion capture datasets. MotionScript serves as both a descriptive tool and a training resource for text-to-motion models, enabling the synthesis of highly realistic and diverse human motions from text. By augmenting motion datasets with MotionScript captions, we demonstrate significant improvements in out-of-distribution motion generation, allowing large language models (LLMs) to generate motions that extend beyond existing data. Additionally, MotionScript opens new applications in animation, virtual human simulation, and robotics, providing an interpretable bridge between intuitive descriptions and motion synthesis. To the best of our knowledge, this is the first attempt to systematically translate 3D motion into structured natural language without requiring training data.
Code (0)
등록된 구현이 없습니다.
Tasks
DescriptiveDiversityMotion GenerationMotion SynthesisMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Expressive TTS Driven by Natural Language Prompts Using Few Human Annotations
Expressive text-to-speech (TTS) aims to synthesize speeches with human-like tones, moods, or even artistic attributes. Recent advancements in expressive TTS empower users with the ability to directly control synthesis st…
Language ModelingLanguage ModellingLarge Language ModelRetrieval+2Harmon: Whole-Body Motion Generation of Humanoid Robots from Language Descriptions
Humanoid robots, with their human-like embodiment, have the potential to integrate seamlessly into human environments. Critical to their coexistence and cooperation with humans is the ability to understand natural langua…
Motion GenerationSpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description
Speech-language multi-modal learning presents a significant challenge due to the fine nuanced information inherent in speech styles. Therefore, a large-scale dataset providing elaborate comprehension of speech style is u…
DescriptiveSpeech SynthesisTAGSemantic Regexes: Auto-Interpreting LLM Features with a Structured Language
Automated interpretability aims to translate large language model (LLM) features into human understandable descriptions. However, natural language feature descriptions can be vague, inconsistent, and require manual relab…
Sketchforme: Composing Sketched Scenes from Text Descriptions for Interactive Applications
Sketching and natural languages are effective communication media for interactive applications. We introduce Sketchforme, the first neural-network-based system that can generate sketches based on text descriptions specif…