paper-with-me

홈 › Papers

Learning to Generate Human-Human-Object Interactions from Textual Descriptions

2025-11-25 · Jeonghyeon Na, Sangwon Baik, Inhee Lee, Junyoung Lee, Hanbyul Joo arxiv

The way humans interact with each other, including interpersonal distances, spatial configuration, and motion, varies significantly across different situations. To enable machines to understand such complex, context-dependent behaviors, it is essential to model multiple people in relation to the surrounding scene context. In this paper, we present a novel research problem to model the correlations between two people engaged in a shared interaction involving an object. We refer to this formulation as Human-Human-Object Interactions (HHOIs). To overcome the lack of dedicated datasets for HHOIs, we present a newly captured HHOIs dataset and a method to synthesize HHOI data by leveraging image generative models. As an intermediary, we obtain individual human-object interaction (HOIs) and human-human interaction (HHIs) from the HHOIs, and with these data, we train an text-to-HOI and text-to-HHI model using score-based diffusion model. Finally, we present a unified generative framework that integrates the two individual model, capable of synthesizing complete HHOIs in a single advanced sampling process. Our method extends HHOI generation to multi-human settings, enabling interactions involving more than two individuals. Experimental results show that our method generates realistic HHOIs conditioned on textual descriptions, outperforming previous approaches that focus only on single-human HOIs. Furthermore, we introduce multi-human motion generation involving objects as an application of our framework.

📄 PDF Abstract BibTeX arXiv:2511.20446

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models

2023-12-11 · Xiaogang Peng, Yiming Xie, Zizhao Wu, Varun Jampani 외

We address the problem of generating realistic 3D human-object interactions (HOIs) driven by textual prompts. To this end, we take a modular design and decompose the complex task into simpler sub-tasks. We first develop …

Human-Object Interaction DetectionMotion GenerationObject

Controllable Human-Object Interaction Synthesis

2023-12-06 · Jiaman Li, Alexander Clegg, Roozbeh Mottaghi, Jiajun Wu 외

Synthesizing semantic-aware, long-horizon, human-object interaction is critical to simulate realistic human behaviors. In this work, we address the challenging problem of generating synchronized object motion and human m…

Human-Object Interaction DetectionObject

Contextualized Representation Learning for Effective Human-Object Interaction Detection

2025-09-16 · Zhehao Li, Yucheng Qian, Chong Wang, Yinghao Lu 외 arxiv

Human-Object Interaction (HOI) detection aims to simultaneously localize human-object pairs and recognize their interactions. While recent two-stage approaches have made significant progress, they still face challenges d…

Human-Object Interaction DetectionRepresentation Learning

Human-Object Interaction from Human-Level Instructions

2024-06-25 · Zhen Wu, Jiaman Li, Pei Xu, C. Karen Liu

Intelligent agents must autonomously interact with the environments to perform daily tasks based on human-level instructions. They need a foundational understanding of the world to accurately interpret these instructions…

Common Sense ReasoningHuman-Object Interaction DetectionLanguage ModellingLarge Language Model+3

ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video Generation

2024-12-24 · Hongjie Li, Hong-Xing Yu, Jiaman Li, Jiajun Wu

Human-scene interaction (HSI) generation is crucial for applications in embodied AI, virtual reality, and robotics. Yet, existing methods cannot synthesize interactions in unseen environments such as in-the-wild scenes o…

Human-Object Interaction DetectionVideo Generation