paper-with-me

Papers

HOIAnimator: Generating Text-prompt Human-object Animations using Novel Perceptive Diffusion Models

2024-01-01 · CVPR 2024 1 · Wenfeng Song, Xinyu Zhang, Shuai Li, Yang Gao, Aimin Hao, Xia Hou, Chenglizhao Chen, Ning li, Hong Qin

To date the quest to rapidly and effectively produce human-object interaction (HOI) animations directly from textual descriptions stands at the forefront of computer vision research. The underlying challenge demands both a discriminating interpretation of language and a comprehensive physics-centric model supporting real-world dynamics. To ameliorate this paper advocates HOIAnimator a novel and interactive diffusion model with perception ability and also ingeniously crafted to revolutionize the animation of complex interactions from linguistic narratives. The effectiveness of our model is anchored in two ground-breaking innovations: (1) Our Perceptive Diffusion Models (PDM) brings together two types of models: one focused on human movements and the other on objects. This combination allows for animations where humans and objects move in concert with each other making the overall motion more realistic. Additionally we propose a Perceptive Message Passing (PMP) mechanism to enhance the communication bridging the two models ensuring that the animations are smooth and unified; (2) We devise an Interaction Contact Field (ICF) a sophisticated model that implicitly captures the essence of HOIs. Beyond mere predictive contact points the ICF assesses the proximity of human and object to their respective environment informed by a probabilistic distribution of interactions learned throughout the denoising phase. Our comprehensive evaluation showcases HOIanimator's superior ability to produce dynamic context-aware animations that surpass existing benchmarks in text-driven animation synthesis.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingHuman-Object Interaction Detection

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance Guidance

2025-06-08 · Lei LI, Angela Dai

We present HOI-PAGE, a new approach to synthesizing 4D human-object interactions (HOIs) from text prompts in a zero-shot fashion, driven by part-level affordance reasoning. In contrast to prior works that focus on global…

Human-Object Interaction DetectionHuman-Object Interaction GenerationObject

MuLan: Multimodal-LLM Agent for Progressive and Interactive Multi-Object Diffusion

2024-02-20 · Sen Li, Ruochen Wang, Cho-Jui Hsieh, Minhao Cheng 외

Existing text-to-image models still struggle to generate images of multiple objects, especially in handling their spatial positions, relative sizes, overlapping, and attribute bindings. To efficiently address these chall…

AttributeLanguage ModelingLanguage ModellingLarge Language Model+1

HIMO: A New Benchmark for Full-Body Human Interacting with Multiple Objects

2024-07-17 · Xintao Lv, Liang Xu, Yichao Yan, Xin Jin 외

Generating human-object interactions (HOIs) is critical with the tremendous advances of digital avatars. Existing datasets are typically limited to humans interacting with a single object while neglecting the ubiquitous …

BenchmarkingHuman-Object Interaction DetectionObject

Hoi3DGen: Generating High-Quality Human-Object-Interactions in 3D

2026-03-12 · Agniv Sharma, Xianghui Xie, Tom Fischer, Eddy Ilg 외 arxiv

Modeling and generating 3D human-object interactions from text is crucial for applications in AR, XR, and gaming. Existing approaches often rely on score distillation from text-to-image models, but their results suffer f…

3D Generation

HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models

2023-12-11 · Xiaogang Peng, Yiming Xie, Zizhao Wu, Varun Jampani 외

We address the problem of generating realistic 3D human-object interactions (HOIs) driven by textual prompts. To this end, we take a modular design and decompose the complex task into simpler sub-tasks. We first develop …

Human-Object Interaction DetectionMotion GenerationObject