paper-with-me

Papers

DiMo: Discrete Diffusion Modeling for Motion Generation and Understanding

2026-02-04 · Ning Zhang, Zhengyu Li, Kwong Weng Loh, Mingxi Xu, Qi Wang, Zhengyu Wen, Xiaoyu He, Wei Zhao, Kehong Gong, Mingyuan Zhang arxiv

Prior masked modeling motion generation methods predominantly study text-to-motion. We present DiMo, a discrete diffusion-style framework, which extends masked modeling to bidirectional text--motion understanding and generation. Unlike GPT-style autoregressive approaches that tokenize motion and decode sequentially, DiMo performs iterative masked token refinement, unifying Text-to-Motion (T2M), Motion-to-Text (M2T), and text-free Motion-to-Motion (M2M) within a single model. This decoding paradigm naturally enables a quality-latency trade-off at inference via the number of refinement steps. We further improve motion token fidelity with residual vector quantization (RVQ) and enhance alignment and controllability with Group Relative Policy Optimization (GRPO). Experiments on HumanML3D and KIT-ML show strong motion quality and competitive bidirectional understanding under a unified framework. In addition, we demonstrate model ability in text-free motion completion, text-guided motion prediction and motion caption correction without architectural change. Additional qualitative results are available on our project page: https://animotionlab.github.io/DiMo/.

📄 PDF Abstract BibTeX arXiv:2602.04188

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding

2025-10-07 · Yi Xin, Qi Qin, Siqi Luo, Kaiwen Zhu 외 arxiv

We introduce Lumina-DiMOO, an open-source foundational model for seamless multi-modal generation and understanding. Lumina-DiMOO sets itself apart from prior unified models by utilizing a fully discrete diffusion modelin…

Text-to-Image GenerationImage InpaintingImage Editing

Rarity-Aware Discrete Diffusion with Spatially Consistent Decoding for Photo-Realistic Image Super-Resolution

2026-07-20 · Ao Li, Yapeng Du, Yi Xin, Lei Zhu 외 arxiv

Continuous diffusion models have become the dominant paradigm for photo-realistic image Super-Resolution (SR), but they typically formulate reconstruction as continuous signal-level denoising and incorporate semantic pri…

Image Super-Resolution

DC-Motion: Decoupling Structure and Details via Discrete-Continuous Tokens for Human Motion Generation

2026-05-28 · Hequan Wang, Xuean Chen, Jiaxu Zhang, Zhengbo Zhang 외 arxiv

Text-to-motion generation requires modeling both global action structure and fine-grained motion dynamics from natural language. Existing approaches typically rely on either continuous diffusion models or vector-quantize…

DIMO: Diverse 3D Motion Generation for Arbitrary Objects

2025-11-10 · Linzhan Mou, Jiahui Lei, Chen Wang, Lingjie Liu 외 arxiv

We present DIMO, a generative approach capable of generating diverse 3D motions for arbitrary objects from a single image. The core idea of our work is to leverage the rich priors in well-trained video models to extract …

Domain-conditioned and Temporal-guided Diffusion Modeling for Accelerated Dynamic MRI Reconstruction

2025-01-16 · Liping Zhang, Iris Yuwen Zhou, Sydney B. Montesi, Li Feng 외

Purpose: To propose a domain-conditioned and temporal-guided diffusion modeling method, termed dynamic Diffusion Modeling (dDiMo), for accelerated dynamic MRI reconstruction, enabling diffusion process to characterize sp…

MRI Reconstruction