paper-with-me

홈 › Papers

M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation

2024-05-25 · Mingshuang Luo, Ruibing Hou, Zhuo Li, Hong Chang, Zimo Liu, YaoWei Wang, Shiguang Shan

This paper presents M$^3$GPT, an advanced $\textbf{M}$ultimodal, $\textbf{M}$ultitask framework for $\textbf{M}$otion comprehension and generation. M$^3$GPT operates on three fundamental principles. The first focuses on creating a unified representation space for various motion-relevant modalities. We employ discrete vector quantization for multimodal conditional signals, such as text, music and motion/dance, enabling seamless integration into a large language model (LLM) with a single vocabulary. The second involves modeling motion generation directly in the raw motion space. This strategy circumvents the information loss associated with a discrete tokenizer, resulting in more detailed and comprehensive motion generation. Third, M$^3$GPT learns to model the connections and synergies among various motion-relevant tasks. Text, the most familiar and well-understood modality for LLMs, is utilized as a bridge to establish connections between different motion tasks, facilitating mutual reinforcement. To our knowledge, M$^3$GPT is the first model capable of comprehending and generating motions based on multiple signals. Extensive experiments highlight M$^3$GPT's superior performance across various motion-relevant tasks and its powerful zero-shot generalization capabilities for extremely challenging tasks. Project page: \url{https://github.com/luomingshuang/M3GPT}.

📄 PDF Abstract BibTeX arXiv:2405.16273

Code (1)

luomingshuang/m3gpt 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingLarge Language ModelMotion GenerationQuantizationZero-shot Generalization

Similar Papers 제목 키워드 기반

MotionLLaMA: A Unified Framework for Motion Synthesis and Comprehension

2024-11-26 · Zeyu Ling, Bo Han, Shiyang Li, Hongdeng Shen 외

This paper introduces MotionLLaMA, a unified framework for motion synthesis and comprehension, along with a novel full-body motion tokenizer called the HoMi Tokenizer. MotionLLaMA is developed based on three core princip…

Language ModelingLanguage ModellingLarge Language ModelMotion Synthesis+1

COMMA-DEER: COmmon-sense Aware Multimodal Multitask Approach for Detection of Emotion and Emotional Reasoning in Conversations

2022-10-01 · COLING 2022 10 · Soumitra Ghosh, Gopendra Vikram Singh, Asif Ekbal, Pushpak Bhattacharyya

Mental health is a critical component of the United Nations’ Sustainable Development Goals (SDGs), particularly Goal 3, which aims to provide “good health and well-being”. The present mental health treatment gap is exace…

Common Sense Reasoning

A Sentiment and Emotion Aware Multimodal Multiparty Humor Recognition in Multilingual Conversational Setting

2022-10-01 · COLING 2022 10 · Dushyant Singh Chauhan, Gopendra Vikram Singh, Aseem Arora, Asif Ekbal 외

In this paper, we hypothesize that humor is closely related to sentiment and emotions. Also, due to the tremendous growth in multilingual content, there is a great demand for building models and systems that support mult…

Humor Detection

A Multimodal-Multitask Framework with Cross-modal Relation and Hierarchical Interactive Attention for Semantic Comprehension

2025-08-22 · Mohammad Zia Ur Rehman, Devraj Raghuvanshi, Umang Jain, Shubhi Bansal 외 arxiv

A major challenge in multimodal learning is the presence of noise within individual modalities. This noise inherently affects the resulting multimodal representations, especially when these representations are obtained t…

An Emoji-aware Multitask Framework for Multimodal Sarcasm Detection

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Sarcasm is a case of implicit emotion and needs additional information like context and multimodality for its better detection. But sometimes this additional information also fails to help in sarcasm detection. For examp…

Sarcasm Detection