paper-with-me

홈 › Papers

Human-Aware 3D Scene Generation with Spatially-constrained Diffusion Models

2024-06-26 · Xiaolin Hong, Hongwei Yi, Fazhi He, Qiong Cao

Generating 3D scenes from human motion sequences supports numerous applications, including virtual reality and architectural design. However, previous auto-regression-based human-aware 3D scene generation methods have struggled to accurately capture the joint distribution of multiple objects and input humans, often resulting in overlapping object generation in the same space. To address this limitation, we explore the potential of diffusion models that simultaneously consider all input humans and the floor plan to generate plausible 3D scenes. Our approach not only satisfies all input human interactions but also adheres to spatial constraints with the floor plan. Furthermore, we introduce two spatial collision guidance mechanisms: human-object collision avoidance and object-room boundary constraints. These mechanisms help avoid generating scenes that conflict with human motions while respecting layout constraints. To enhance the diversity and accuracy of human-guided scene generation, we have developed an automated pipeline that improves the variety and plausibility of human-object interactions in the existing 3D FRONT HUMAN dataset. Extensive experiments on both synthetic and real-world datasets demonstrate that our framework can generate more natural and plausible 3D scenes with precise human-scene interactions, while significantly reducing human-object collisions compared to previous state-of-the-art methods. Our code and data will be made publicly available upon publication of this work.

📄 PDF Abstract BibTeX arXiv:2406.18159

Code (0)

등록된 구현이 없습니다.

Tasks

Collision AvoidanceHuman-Object Interaction DetectionObjectScene Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Audio Visual Scene-Aware Dialog Generation with Transformer-based Video Representations

2022-02-21 · Yoshihiro Yamazaki, Shota Orihashi, Ryo Masumura, Mihiro Uchida 외

There have been many attempts to build multimodal dialog systems that can respond to a question about given audio-visual information, and the representative task for such systems is the Audio Visual Scene-Aware Dialog (A…

Answer GenerationVideo Understanding

GenSpace: Benchmarking Spatially-Aware Image Generation

2025-05-30 · Zehan Wang, Jiayang Xu, Ziang Zhang, Tianyu Pan 외

Humans can intuitively compose and arrange scenes in the 3D space for photography. However, can advanced AI image generators plan scenes with similar 3D spatial awareness when creating images from text or image prompts? …

BenchmarkingImage Generation

Scene and Human in One World: Reconstruction in a Feedforward Pass

2026-06-26 · Boao Shi, Qiao Feng, Yiming Huang, Lingjie Liu arxiv

Reconstructing humans in dynamic scenes from moving monocular cameras remains challenging due to scale ambiguity, human-scene misalignment, and occlusion interference. Rather than treating human mesh recovery and scene r…

Human Mesh Recovery

R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation

2025-10-09 · Xiuwei Xu, Angyuan Ma, Hankun Li, Bingyao Yu 외 arxiv

Towards the aim of generalized robotic manipulation, spatial generalization is the most fundamental capability that requires the policy to work robustly under different spatial distribution of objects, environment and ag…

DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis

2022-12-22 · CVPR 2023 1 · Yinghao Xu, Menglei Chai, Zifan Shi, Sida Peng 외

Existing 3D-aware image synthesis approaches mainly focus on generating a single canonical object and show limited capacity in composing a complex scene containing a variety of objects. This work presents DisCoScene: a 3…

3D-Aware Image SynthesisImage GenerationObject