paper-with-me

홈 › Papers

Marrying Text-to-Motion Generation with Skeleton-Based Action Recognition

2026-04-18 · Jidong Kuang, Hongsong Wang, Jie Gui arxiv

Human action recognition and motion generation are two active research problems in human-centric computer vision, both aiming to align motion with textual semantics. However, most existing works study these two problems separately, without uncovering the links between them, namely that motion generation requires semantic comprehension. This work investigates unified action recognition and motion generation by leveraging skeleton coordinates for both motion understanding and generation. We propose Coordinates-based Autoregressive Motion Diffusion (CoAMD), which synthesizes motion in a coarse-to-fine manner. As a core component of CoAMD, we design a Multi-modal Action Recognizer (MAR) that provides gradient-based semantic guidance for motion generation. Furthermore, we establish a rigorous benchmark by evaluating baselines on absolute coordinates. Our model can be applied to four important tasks, including skeleton-based action recognition, text-to-motion generation, text-motion retrieval, and motion editing. Extensive experiments on 13 benchmarks across these tasks demonstrate that our approach achieves state-of-the-art performance, highlighting its effectiveness and versatility for human motion modeling. Code is available at https://github.com/jidongkuang/CoAMD.

📄 PDF Abstract BibTeX arXiv:2604.17090

Code (0)

등록된 구현이 없습니다.

Tasks

Action Recognition

Similar Papers 제목 키워드 기반

SkeletonGaussian: Editable 4D Generation through Gaussian Skeletonization

2026-02-04 · Lifan Wu, Ruijie Zhu, Yubo Ai, Tianzhu Zhang arxiv

4D generation has made remarkable progress in synthesizing dynamic 3D objects from input text, images, or videos. However, existing methods often represent motion as an implicit deformation field, which limits direct con…

Topology-Agnostic Animal Motion Generation from Text Prompt

2025-12-11 · Keyi Chen, Mingze Sun, Zhenyu Liu, Zhangquan Chen 외 arxiv

Motion generation is fundamental to computer animation and widely used across entertainment, robotics, and virtual environments. While recent methods achieve impressive results, most rely on fixed skeletal templates, whi…

Style Transfer

Shape Conditioned Human Motion Generation with Diffusion Model

2024-05-10 · Kebing Xue, Hyewon Seo

Human motion synthesis is an important task in computer graphics and computer vision. While focusing on various conditioning signals such as text, action class, or audio to guide the generation process, most existing met…

modelMotion GenerationMotion Synthesis

Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades

2026-03-09 · Ashkan Taghipour, Morteza Ghahremani, Zinuo Li, Hamid Laga 외 arxiv

Generating videos of complex human motions such as flips, cartwheels, and martial arts remains challenging for current video diffusion models. Text-only conditioning is temporally ambiguous for fine-grained motion contro…

Video Generation

LAC: Latent Action Composition for Skeleton-based Action Segmentation

2023-08-28 · Di Yang, Yaohui Wang, Antitza Dantcheva, Quan Kong 외

Skeleton-based action segmentation requires recognizing composable actions in untrimmed videos. Current approaches decouple this problem by first extracting local visual features from skeleton sequences and then processi…

Action SegmentationContrastive LearningSegmentationSkeleton Based Action Segmentation+1