paper-with-me

Papers

MDMP: Multi-modal Diffusion for supervised Motion Predictions with uncertainty

2024-10-04 · Leo Bringer, Joey Wilson, Kira Barton, Maani Ghaffari

This paper introduces a Multi-modal Diffusion model for Motion Prediction (MDMP) that integrates and synchronizes skeletal data and textual descriptions of actions to generate refined long-term motion predictions with quantifiable uncertainty. Existing methods for motion forecasting or motion generation rely solely on either prior motions or text prompts, facing limitations with precision or control, particularly over extended durations. The multi-modal nature of our approach enhances the contextual understanding of human motion, while our graph-based transformer framework effectively capture both spatial and temporal motion dynamics. As a result, our model consistently outperforms existing generative techniques in accurately predicting long-term motions. Additionally, by leveraging diffusion models' ability to capture different modes of prediction, we estimate uncertainty, significantly improving spatial awareness in human-robot interactions by incorporating zones of presence with varying confidence levels for each body joint.

📄 PDF Abstract BibTeX arXiv:2410.03860

Code (1)

leob03/mdmp 공식 구현 pytorch

Tasks

Motion ForecastingMotion Generationmotion prediction

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Incomplete Multimodality-Diffused Emotion Recognition

2023-09-21 · NeurIPS 2023 11

Human multimodal emotion recognition (MER) aims to perceive and understand human emotions via various heterogeneous modalities, such as language, vision, and acoustic. Compared with unimodality, the complementary informa…

EmoDiffTalk:Emotion-aware Diffusion for Editable 3D Gaussian Talking Head

2025-11-30 · Chang Liu, Tianjiao Jing, Chengcheng Ma, Xuanqi Zhou 외 arxiv

Recent photo-realistic 3D talking head via 3D Gaussian Splatting still has significant shortcoming in emotional expression manipulation, especially for fine-grained and expansive dynamics emotional editing using multi-mo…

GSDNet: Revisiting Incomplete Multimodal-Diffusion from Graph Spectrum Perspective for Conversation Emotion Recognition

2025-06-14 · Yuntao Shou, Jun Yao, Tao Meng, Wei Ai 외

Multimodal emotion recognition in conversations (MERC) aims to infer the speaker's emotional state by analyzing utterance information from multiple sources (i.e., video, audio, and text). Compared with unimodality, a mor…

Emotion RecognitionModality completionMultimodal Emotion Recognition

Unlocking Pretrained LLMs for Motion-Related Multimodal Generation: A Fine-Tuning Approach to Unify Diffusion and Next-Token Prediction

2025-03-08 · Shinichi Tanaka, Zhao Wang, Yoichi Kato, Jun Ohya

In this paper, we propose a unified framework that leverages a single pretrained LLM for Motion-related Multimodal Generation, referred to as MoMug. MoMug integrates diffusion-based continuous motion generation with the …

Motion GenerationMotion Synthesismultimodal generation

Moaw: Unleashing Motion Awareness for Video Diffusion Models

2026-01-19 · Tianqi Zhang, Ziyi Wang, Wenzhao Zheng, Weiliang Chen 외 arxiv

Video diffusion models, trained on large-scale datasets, naturally capture correspondences of shared features across frames. Recent works have exploited this property for tasks such as optical flow prediction and trackin…

Video Generation