DiffusionPhase: Motion Diffusion in Frequency Domain
In this study, we introduce a learning-based method for generating high-quality human motion sequences from text descriptions (e.g., ``A person walks forward"). Existing techniques struggle with motion diversity and smooth transitions in generating arbitrary-length motion sequences, due to limited text-to-motion datasets and the pose representations used that often lack expressiveness or compactness. To address these issues, we propose the first method for text-conditioned human motion generation in the frequency domain of motions. We develop a network encoder that converts the motion space into a compact yet expressive parameterized phase space with high-frequency details encoded, capturing the local periodicity of motions in time and space with high accuracy. We also introduce a conditional diffusion model for predicting periodic motion parameters based on text descriptions and a start pose, efficiently achieving smooth transitions between motion sequences associated with different text descriptions. Experiments demonstrate that our approach outperforms current methods in generating a broader variety of high-quality motions, and synthesizing long sequences with natural transitions.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversityMotion GenerationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Time Series Diffusion in the Frequency Domain
Fourier analysis has been an instrumental tool in the development of signal processing. This leads us to wonder whether this framework could similarly benefit generative modelling. In this paper, we explore this question…
DenoisingInductive BiasTime SeriesFree-T2M: Frequency Enhanced Text-to-Motion Diffusion Model With Consistency Loss
Rapid progress in text-to-motion generation has been largely driven by diffusion models. However, existing methods focus solely on temporal modeling, thereby overlooking frequency-domain analysis. We identify two key pha…
DenoisingMotion GenerationMotion SynthesisFTMoMamba: Motion Generation with Frequency and Text State Space Models
Diffusion models achieve impressive performance in human motion generation. However, current approaches typically ignore the significance of frequency-domain information in capturing fine-grained motions within the laten…
Motion GenerationSentenceState Space ModelsMotion Artifact Removal in Pixel-Frequency Domain via Alternate Masks and Diffusion Model
Motion artifacts present in magnetic resonance imaging (MRI) can seriously interfere with clinical diagnosis. Removing motion artifacts is a straightforward solution and has been extensively studied. However, paired data…
Frequency-Guided Diffusion Model with Perturbation Training for Skeleton-Based Video Anomaly Detection
Video anomaly detection (VAD) is a vital yet complex open-set task in computer vision, commonly tackled through reconstruction-based methods. However, these methods struggle with two key limitations: (1) insufficient rob…
Anomaly DetectionVideo Anomaly Detection