paper-with-me

Papers

InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object Interaction

2024-03-28 · Sirui Xu, Ziyin Wang, Yu-Xiong Wang, Liang-Yan Gui

Text-conditioned human motion generation has experienced significant advancements with diffusion models trained on extensive motion capture data and corresponding textual annotations. However, extending such success to 3D dynamic human-object interaction (HOI) generation faces notable challenges, primarily due to the lack of large-scale interaction data and comprehensive descriptions that align with these interactions. This paper takes the initiative and showcases the potential of generating human-object interactions without direct training on text-interaction pair data. Our key insight in achieving this is that interaction semantics and dynamics can be decoupled. Being unable to learn interaction semantics through supervised training, we instead leverage pre-trained large models, synergizing knowledge from a large language model and a text-to-motion model. While such knowledge offers high-level control over interaction semantics, it cannot grasp the intricacies of low-level interaction dynamics. To overcome this issue, we further introduce a world model designed to comprehend simple physics, modeling how human actions influence object motion. By integrating these components, our novel framework, InterDreamer, is able to generate text-aligned 3D HOI sequences in a zero-shot manner. We apply InterDreamer to the BEHAVE and CHAIRS datasets, and our comprehensive experimental analysis demonstrates its capability to generate realistic and coherent interaction sequences that seamlessly align with the text directives.

📄 PDF Abstract BibTeX arXiv:2403.19652

Code (0)

등록된 구현이 없습니다.

Tasks

Human-Object Interaction DetectionLanguage ModellingLarge Language ModelMotion GenerationText to 3D

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Dynamic Strategy Chain: Dynamic Zero-Shot CoT for Long Mental Health Support Generation

2023-08-21 · Qi Chen, Dexi Liu

Long counseling Text Generation for Mental health support (LTGM), an innovative and challenging task, aims to provide help-seekers with mental health support through a comprehensive and more acceptable response. The comb…

Text Generation

ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video Generation

2024-12-24 · Hongjie Li, Hong-Xing Yu, Jiaman Li, Jiajun Wu

Human-scene interaction (HSI) generation is crucial for applications in embodied AI, virtual reality, and robotics. Yet, existing methods cannot synthesize interactions in unseen environments such as in-the-wild scenes o…

Human-Object Interaction DetectionVideo Generation

Alternative Semantic Representations for Zero-Shot Human Action Recognition

2017-06-28 · Qian Wang, Ke Chen

A proper semantic representation for encoding side information is key to the success of zero-shot learning. In this paper, we explore two alternative semantic representations especially for zero-shot human action recogni…

Action RecognitionTemporal Action LocalizationZero-Shot Action RecognitionZero-Shot Learning

Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking

2026-06-02 · Zekun Qi, Xuchuan Chen, Dairu Liu, Chenghuai Lin 외 arxiv

We introduce Humanoid-GPT, a GPT-style Transformer with causal attention trained on a billion-scale motion corpus for whole-body control. Unlike prior shallow MLP trackers constrained by scarce data and an agility-genera…

Zero-shot Generalization

EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Expression Recognition

2023-10-25 · Niki Maria Foteinopoulou, Ioannis Patras

Facial Expression Recognition (FER) is a crucial task in affective computing, but its conventional focus on the seven basic emotions limits its applicability to the complex and expanding emotional spectrum. To address th…

Facial Expression RecognitionFacial Expression Recognition (FER)Language Modellingzero-shot-classification+2