paper-with-me

홈 › Papers

DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding

2024-11-29 · Jungbin Cho, Junwan Kim, Jisoo Kim, Minseo Kim, Mingu Kang, Sungeun Hong, Tae-Hyun Oh, Youngjae Yu

Human motion is inherently continuous and dynamic, posing significant challenges for generative models. While discrete generation methods are widely used, they suffer from limited expressiveness and frame-wise noise artifacts. In contrast, continuous approaches produce smoother, more natural motion but often struggle to adhere to conditioning signals due to high-dimensional complexity and limited training data. To resolve this discord between discrete and continuous representations, we introduce DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding, a novel method that leverages rectified flow to decode discrete motion tokens in the continuous, raw motion space. Our core idea is to frame token decoding as a conditional generation task, ensuring that DisCoRD captures fine-grained dynamics and achieves smoother, more natural motions. Compatible with any discrete-based framework, our method enhances naturalness without compromising faithfulness to the conditioning signals on diverse settings. Extensive evaluations Our project page is available at: https://whwjdqls.github.io/discord.github.io/.

📄 PDF Abstract BibTeX arXiv:2411.19527

Code (0)

등록된 구현이 없습니다.

Tasks

Motion SynthesisQuantization

Similar Papers 제목 키워드 기반

From Diffusion to Flow: Efficient Motion Generation in MotionGPT3

2026-03-23 · Jaymin Ban, JiHong Jeon, SangYeop Jeong arxiv

Recent text-driven motion generation methods span both discrete token-based approaches and continuous-latent formulations. MotionGPT3 exemplifies the latter paradigm, combining a learned continuous motion latent space wi…

Audio Generation

DC-Motion: Decoupling Structure and Details via Discrete-Continuous Tokens for Human Motion Generation

2026-05-28 · Hequan Wang, Xuean Chen, Jiaxu Zhang, Zhengbo Zhang 외 arxiv

Text-to-motion generation requires modeling both global action structure and fine-grained motion dynamics from natural language. Existing approaches typically rely on either continuous diffusion models or vector-quantize…

Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions

2025-04-04 · CVPR 2025 1 · Ting-Hsuan Liao, Yi Zhou, Yu Shen, Chun-Hao Paul Huang 외

We explore how body shapes influence human motion synthesis, an aspect often overlooked in existing text-to-motion generation methods due to the ease of learning a homogenized, canonical body shape. However, this homogen…

Language ModelingLanguage ModellingMotion GenerationMotion Synthesis+1

Recovering Performance in Speech Emotion Recognition from Discrete Tokens via Multi-Layer Fusion and Paralinguistic Feature Integration

2026-01-23 · Esther Sun, Abinay Reddy Naini, Carlos Busso arxiv

Discrete speech tokens offer significant advantages for storage and language model integration, but their application in speech emotion recognition (SER) is limited by paralinguistic information loss during quantization.…

Speech Emotion Recognition

Soft Tokens, Hard Truths

2025-09-23 · Natasha Butt, Ariel Kwiatkowski, Ismail Labiad, Julia Kempe 외 arxiv

The use of continuous instead of discrete tokens during the Chain-of-Thought (CoT) phase of reasoning LLMs has garnered attention recently, based on the intuition that a continuous mixture of discrete tokens could simula…

Reinforcement Learning