paper-with-me

홈 › Papers

in2IN: Leveraging individual Information to Generate Human INteractions

2024-04-15 · Pablo Ruiz Ponce, German Barquero, Cristina Palmero, Sergio Escalera, Jose Garcia-Rodriguez

Generating human-human motion interactions conditioned on textual descriptions is a very useful application in many areas such as robotics, gaming, animation, and the metaverse. Alongside this utility also comes a great difficulty in modeling the highly dimensional inter-personal dynamics. In addition, properly capturing the intra-personal diversity of interactions has a lot of challenges. Current methods generate interactions with limited diversity of intra-person dynamics due to the limitations of the available datasets and conditioning strategies. For this, we introduce in2IN, a novel diffusion model for human-human motion generation which is conditioned not only on the textual description of the overall interaction but also on the individual descriptions of the actions performed by each person involved in the interaction. To train this model, we use a large language model to extend the InterHuman dataset with individual descriptions. As a result, in2IN achieves state-of-the-art performance in the InterHuman dataset. Furthermore, in order to increase the intra-personal diversity on the existing interaction datasets, we propose DualMDM, a model composition technique that combines the motions generated with in2IN and the motions generated by a single-person motion prior pre-trained on HumanML3D. As a result, DualMDM generates motions with higher individual diversity and improves control over the intra-person dynamics while maintaining inter-personal coherence.

📄 PDF Abstract BibTeX arXiv:2404.09988

Code (1)

pabloruizponce/in2IN 공식 구현 pytorch

Tasks

DiversityLanguage ModellingLarge Language ModelMotion GenerationMotion Synthesis

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Learning to Generate Human-Human-Object Interactions from Textual Descriptions

2025-11-25 · Jeonghyeon Na, Sangwon Baik, Inhee Lee, Junyoung Lee 외 arxiv

The way humans interact with each other, including interpersonal distances, spatial configuration, and motion, varies significantly across different situations. To enable machines to understand such complex, context-depe…

Human Interaction-Aware 3D Reconstruction from a Single Image

2026-04-07 · Gwanghyun Kim, Junghun James Kim, Suh Yoon Jeon, Jason Park 외 arxiv

Reconstructing textured 3D human models from a single image is fundamental for AR/VR and digital human applications. However, existing methods mostly focus on single individuals and thus fail in multi-human scenes, where…

3D Reconstruction

SocialGen: Modeling Multi-Human Social Interaction with Language Models

2025-03-28 · Heng Yu, Juze Zhang, Changan Chen, Tiange Xiang 외

Human interactions in everyday life are inherently social, involving engagements with diverse individuals across various contexts. Modeling these social interactions is fundamental to a wide range of real-world applicati…

Towards Better Adversarial Synthesis of Human Images from Text

2021-07-05 · Rania Briq, Pratika Kochar, Juergen Gall

This paper proposes an approach that generates multiple 3D human meshes from text. The human shapes are represented by 3D meshes based on the SMPL model. The model's performance is evaluated on the COCO dataset, which co…

Image Generation

Multi-Person Interaction Generation from Two-Person Motion Priors

2025-05-23 · Wenning Xu, Shiyu Fan, Paul Henderson, Edmond S. L. Ho

Generating realistic human motion with high-level controls is a crucial task for social understanding, robotics, and animation. With high-quality MOCAP data becoming more available recently, a wide range of data-driven a…

Motion Generation