paper-with-me

홈 › Papers

Multi-granular body modeling with Redundancy-Free Spatiotemporal Fusion for Text-Driven Motion Generation

2025-03-10 · Xingzu Zhan, Chen Xie, Honghang Chen, Haoran Sun, Xiaochun Mai

Text-to-motion generation sits at the intersection of multimodal learning and computer graphics and is gaining momentum because it can simplify content creation for games, animation, robotics and virtual reality. Most current methods stack spatial and temporal features in a straightforward way, which adds redundancy and still misses subtle joint-level cues. We introduce HiSTF Mamba, a framework with three parts: Dual-Spatial Mamba, Bi-Temporal Mamba and a Dynamic Spatiotemporal Fusion Module (DSFM). The Dual-Spatial module runs part-based and whole-body models in parallel, capturing both overall coordination and fine-grained joint motion. The Bi-Temporal module scans sequences forward and backward to encode short-term details and long-term dependencies. DSFM removes redundant temporal information, extracts complementary cues and fuses them with spatial features to build a richer spatiotemporal representation. Experiments on the HumanML3D benchmark show that HiSTF Mamba performs well across several metrics, achieving high fidelity and tight semantic alignment between text and motion.

📄 PDF Abstract BibTeX arXiv:2503.06897

Code (0)

등록된 구현이 없습니다.

Tasks

MambaMotion Generation

Methods 이 논문이 사용한 방법론

Mamba Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module.…

Similar Papers 제목 키워드 기반

GaitMM: Multi-Granularity Motion Sequence Learning for Gait Recognition

2022-09-18 · Lei Wang, Bo Liu, Bincheng Wang, Fuqiang Yu

Gait recognition aims to identify individual-specific walking patterns by observing the different periodic movements of each body part. However, most existing methods treat each part equally and fail to account for the d…

Gait RecognitionMultiview Gait Recognition

M3G: Multi-Granular Gesture Generator for Audio-Driven Full-Body Human Motion Synthesis

2025-05-13 · Zhizhuo Yin, Yuk Hang Tsui, Pan Hui

Generating full-body human gestures encompassing face, body, hands, and global movements from audio is a valuable yet challenging task in virtual avatar creation. Previous systems focused on tokenizing the human gestures…

Gesture GenerationMotion Synthesis

Granular Computing-driven SAM: From Coarse-to-Fine Guidance for Prompt-Free Segmentation

2025-11-24 · Qiyang Yu, Yu Fang, Tianrui Li, Xuemei Cao 외 arxiv

Prompt-free image segmentation aims to generate accurate masks without manual guidance. Typical pre-trained models, notably Segmentation Anything Model (SAM), generate prompts directly at a single granularity level. Howe…

Image Segmentation

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

2025-07-10 · Jeongseok Hyun, Sukjun Hwang, Su Ho Han, Taeoh Kim 외 arxiv

Video large language models (LLMs) achieve strong video understanding by leveraging a large number of spatio-temporal tokens, but suffer from quadratic computational scaling with token count. To address this, we propose …

MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities

2025-04-03 · CVPR 2025 1 · Bizhu Wu, Jinheng Xie, Keming Shen, Zhe Kong 외

Recent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where …

Language ModelingLanguage Modelling