PersonaBooth: Personalized Text-to-Motion Generation
This paper introduces Motion Personalization, a new task that generates personalized motions aligned with text descriptions using several basic motions containing Persona. To support this novel task, we introduce a new large-scale motion dataset called PerMo (PersonaMotion), which captures the unique personas of multiple actors. We also propose a multi-modal finetuning method of a pretrained motion diffusion model called PersonaBooth. PersonaBooth addresses two main challenges: i) A significant distribution gap between the persona-focused PerMo dataset and the pretraining datasets, which lack persona-specific data, and ii) the difficulty of capturing a consistent persona from the motions vary in content (action type). To tackle the dataset distribution gap, we introduce a persona token to accept new persona features and perform multi-modal adaptation for both text and visuals during finetuning. To capture a consistent persona, we incorporate a contrastive learning technique to enhance intra-cohesion among samples with the same persona. Furthermore, we introduce a context-aware fusion mechanism to maximize the integration of persona cues from multiple input motions. PersonaBooth outperforms state-of-the-art motion style transfer methods, establishing a new benchmark for motion personalization.
Code (0)
등록된 구현이 없습니다.
Tasks
Contrastive LearningMotion GenerationMotion Style TransferStyle TransferMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Few-Shot Human Motion Transfer by Personalized Geometry and Texture Modeling
We present a new method for few-shot human motion transfer that achieves realistic human image generation with only a small number of appearance inputs. Despite recent advances in single person motion transfer, prior met…
Appearance TransferImage GenerationDreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models
Emotional talking head generation has attracted growing attention. Previous methods, which are mainly GAN-based, still struggle to consistently produce satisfactory results across diverse emotions and cannot conveniently…
DenoisingTalking Head GenerationSecure & Personalized Music-to-Video Generation via CHARCHA
Music is a deeply personal experience and our aim is to enhance this with a fully-automated pipeline for personalized music video generation. Our work allows listeners to not just be consumers but co-creators in the musi…
RhythmVideo GenerationAnimateLCM: Computation-Efficient Personalized Style Video Generation without Personalized Video Data
This paper introduces an effective method for computation-efficient personalized style video generation without requiring access to any personalized video data. It reduces the necessary generation time of similarly sized…
Conditional Image GenerationDenoisingImage GenerationMotion Generation+1PersonaAnimator: Personalized Motion Transfer from Unconstrained Videos
Recent advances in motion generation show remarkable progress. However, several limitations remain: (1) Existing pose-guided character motion transfer methods merely replicate motion without learning its style characteri…
Style Transfer