paper-with-me

홈 › Papers

FantasyHSI: Video-Generation-Centric 4D Human Synthesis In Any Scene through A Graph-based Multi-Agent Framework

2025-09-01 · Lingzhou Mu, Qiang Wang, Fan Jiang, Mengchao Wang, Yaqi Fan, Mu Xu, Kai Zhang arxiv

Human-Scene Interaction (HSI) seeks to generate realistic human behaviors within complex environments, yet it faces significant challenges in handling long-horizon, high-level tasks and generalizing to unseen scenes. To address these limitations, we introduce FantasyHSI, a novel HSI framework centered on video generation and multi-agent systems that operates without paired data. We model the complex interaction process as a dynamic directed graph, upon which we build a collaborative multi-agent system. This system comprises a scene navigator agent for environmental perception and high-level path planning, and a planning agent that decomposes long-horizon goals into atomic actions. Critically, we introduce a critic agent that establishes a closed-loop feedback mechanism by evaluating the deviation between generated actions and the planned path. This allows for the dynamic correction of trajectory drifts caused by the stochasticity of the generative model, thereby ensuring long-term logical consistency. To enhance the physical realism of the generated motions, we leverage Direct Preference Optimization (DPO) to train the action generator, significantly reducing artifacts such as limb distortion and foot-sliding. Extensive experiments on our custom SceneBench benchmark demonstrate that FantasyHSI significantly outperforms existing methods in terms of generalization, long-horizon task completion, and physical realism. Ours project page: https://fantasy-amap.github.io/fantasy-hsi/

📄 PDF Abstract BibTeX arXiv:2509.01232

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis

2025-11-25 · Mohammad Mahdi, Yuqian Fu, Nedko Savov, Jiancheng Pan 외 arxiv

Foundation video generation models such as WAN 2.2 exhibit strong text- and image-conditioned synthesis abilities but remain constrained to the same-view generation setting. In this work, we introduce Exo2EgoSyn, an adap…

Video Generation

EgoTwin: Dreaming Body and View in First Person

2025-08-18 · Jingqiao Xiu, Fangzhou Hong, Yicong Li, Mengze Li 외 arxiv

While exocentric video synthesis has achieved great progress, egocentric video generation remains largely underexplored, which requires modeling first-person view content along with camera motion patterns induced by the …

Video Generation

HumanForge: A Human-Centric Deepfake Video Benchmark with Multi-Agent Forgery Rationales

2026-07-09 · Wenbo Xu, Zhimin Chen, Xiaojie Liang, Hengrui Liu 외 arxiv

Rapid advancements in video diffusion models and temporal editing tools have enabled the generation of highly realistic human-centric videos, presenting unprecedented challenges to digital content forensics. Existing ben…

Zero-shot Generalization

MV-Performer: Taming Video Diffusion Model for Faithful and Synchronized Multi-view Performer Synthesis

2025-10-08 · Yihao Zhi, Chenghong Li, Hongjie Liao, Xihe Yang 외 arxiv

Recent breakthroughs in video generation, powered by large-scale datasets and diffusion techniques, have shown that video diffusion models can function as implicit 4D novel view synthesizers. Nevertheless, current method…

Monocular Depth EstimationNovel View SynthesisVideo GenerationPoint Clouds

OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation

2026-04-20 · Lei Zhu, Xing Cai, Yingjie Chen, Yiheng Li 외 arxiv

Recent advancements in audio-video joint generation models have demonstrated impressive capabilities in content creation. However, generating high-fidelity human-centric videos in complex, real-world physical scenes rema…

Video Generation