paper-with-me

홈 › Papers

Text-controlled Motion Mamba: Text-Instructed Temporal Grounding of Human Motion

2024-04-17 · Xinghan Wang, Zixi Kang, Yadong Mu

Human motion understanding is a fundamental task with diverse practical applications, facilitated by the availability of large-scale motion capture datasets. Recent studies focus on text-motion tasks, such as text-based motion generation, editing and question answering. In this study, we introduce the novel task of text-based human motion grounding (THMG), aimed at precisely localizing temporal segments corresponding to given textual descriptions within untrimmed motion sequences. Capturing global temporal information is crucial for the THMG task. However, transformer-based models that rely on global temporal self-attention face challenges when handling long untrimmed sequences due to the quadratic computational cost. We address these challenges by proposing Text-controlled Motion Mamba (TM-Mamba), a unified model that integrates temporal global context, language query control, and spatial graph topology with only linear memory cost. The core of the model is a text-controlled selection mechanism which dynamically incorporates global temporal information based on text query. The model is further enhanced to be topology-aware through the integration of relational embeddings. For evaluation, we introduce BABEL-Grounding, the first text-motion dataset that provides detailed textual descriptions of human actions along with their corresponding temporal segments. Extensive evaluations demonstrate the effectiveness of TM-Mamba on BABEL-Grounding.

📄 PDF Abstract BibTeX arXiv:2404.11375

Code (0)

등록된 구현이 없습니다.

Tasks

MambaMotion GenerationQuestion Answering

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Quality-Controlled Multimodal Emotion Recognition in Conversations with Identity-Based Transfer Learning and MAMBA Fusion

2025-11-18 · Zanxu Wang, Homayoon Beigi arxiv

This paper addresses data quality issues in multimodal emotion recognition in conversation (MERC) through systematic quality control and multi-stage transfer learning. We implement a quality control pipeline for MELD and…

Multimodal Emotion RecognitionTransfer LearningFace RecognitionFace Detection

Synthesizing Long-Term Human Motions with Diffusion Models via Coherent Sampling

2023-08-03 · Zhao Yang, Bing Su, Ji-Rong Wen

Text-to-motion generation has gained increasing attention, but most existing methods are limited to generating short-term motions that correspond to a single sentence describing a single action. However, when a text stre…

Motion GenerationSentence

FTMoMamba: Motion Generation with Frequency and Text State Space Models

2024-11-26 · Chengjian Li, Xiangbo Shu, Qiongjie Cui, Yazhou Yao 외

Diffusion models achieve impressive performance in human motion generation. However, current approaches typically ignore the significance of frequency-domain information in capturing fine-grained motions within the laten…

Motion GenerationSentenceState Space Models

KMM: Key Frame Mask Mamba for Extended Motion Generation

2024-11-10 · Zeyu Zhang, Hang Gao, Akide Liu, Qi Chen 외

Human motion generation is a cut-edge area of research in generative computer vision, with promising applications in video creation, game development, and robotic manipulation. The recent Mamba architecture shows promisi…

Contrastive LearningMambaMotion Generation

Multi-granular body modeling with Redundancy-Free Spatiotemporal Fusion for Text-Driven Motion Generation

2025-03-10 · Xingzu Zhan, Chen Xie, Honghang Chen, Haoran Sun 외

Text-to-motion generation sits at the intersection of multimodal learning and computer graphics and is gaining momentum because it can simplify content creation for games, animation, robotics and virtual reality. Most cu…

MambaMotion Generation