paper-with-me

홈 › Papers

Odoriko: A Shape-Aware Multimodal Diffusion Framework for Human Motion

2026-06-19 · Dongseok Shim, Julian Tanke, Kengo Uchida, Christian Simon, Koichi Saito, Takashi Shibuya, Shusuke Takahashi, Yuki Mitsufuji arxiv

Human motion generation has been widely studied across diverse input modalities, text, music, and video, and recent efforts have unified these into single multimodal frameworks. However, while morphological factors such as gender and body shape are known to produce distinct kinematic signatures, no existing unified framework incorporates this into generation, treating all subjects as morphologically equivalent. We present Odoriko, the first unified multimodal motion generation framework that reflects subject bio-morphological information directly in synthesized motion output. Rather than averaging over subject variation, Odoriko generates motion that is consistent with who is moving, not just what they are asked to do, across text, music, and video conditions within a single model. When explicit morphological information is unavailable, Odoriko additionally recovers subject morphology alongside motion, unifying estimation and generation in one framework. Extensive experiments across text-to-motion, music-to-dance, and video-to-motion benchmarks demonstrate that Odoriko matches or exceeds prior specialized models on standard metrics, while enabling morphology-consistent generation that no existing unified framework supports.

📄 PDF Abstract BibTeX arXiv:2606.21135

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sketch2CT: Multimodal Diffusion for Structure-Aware 3D Medical Volume Generation

2026-03-23 · Delin An, Chaoli Wang arxiv

Diffusion probabilistic models have demonstrated significant potential in generating high-quality, realistic medical images, providing a promising solution to the persistent challenge of data scarcity in the medical fiel…

Vitality-Aware Compression for Efficient Image-to-Shape Diffusion Transformers

2026-07-01 · Jaeah Lee, Hyunjin Kim, Jaewoong Cho, Gihyun Kwon arxiv

We propose the first compression approach for image-to-shape Diffusion Transformers (DiTs) that substantially reduces model size while preserving geometric fidelity. Despite remarkable progress in 3D shape generation, la…

Model Compression3D Generation

Variational Shape Inference for Grasp Diffusion on SE(3)

2025-08-24 · S. Talha Bukhari, Kaivalya Agrawal, Zachary Kingston, Aniket Bera arxiv

Grasp synthesis is a fundamental task in robotic manipulation which usually has multiple feasible solutions. Multimodal grasp synthesis seeks to generate diverse sets of stable grasps conditioned on object geometry, maki…

Locally Attentional SDF Diffusion for Controllable 3D Shape Generation

2023-05-08 · Xin-Yang Zheng, Hao Pan, Peng-Shuai Wang, Xin Tong 외

Although the recent rapid evolution of 3D generative neural networks greatly improves 3D shape generation, it is still not convenient for ordinary users to create 3D shapes and control the local geometry of generated sha…

3D Generation3D Shape Generation

DiffusionCom: Structure-Aware Multimodal Diffusion Model for Multimodal Knowledge Graph Completion

2025-04-09 · Wei Huang, Meiyu Liang, Peining Li, Xu Hou 외

Most current MKGC approaches are predominantly based on discriminative models that maximize conditional likelihood. These approaches struggle to efficiently capture the complex connections in real-world knowledge graphs,…

Graph AttentionKnowledge Graph CompletionKnowledge GraphsRepresentation Learning+1