paper-with-me

홈 › Papers

Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes

2026-05-22 · Wenxuan Peng, Bharath Hariharan, Hadar Averbuch-Elor arxiv

Despite recent progress, text-to-image models still struggle to generate semantically diverse and compositionally accurate multi-person interaction scenes, often collapsing to repetitive layouts, stereotypical poses, and poorly grounded interactions. In this work, we bridge this gap by introducing a dual pose-image representation that brings person-centric structural priors into pretrained diffusion transformers. Our model jointly predicts a 2D pose visualization image and its corresponding RGB image, enabling structure and appearance to co-evolve during learning. At its core, a cross-modal alignment scheme binds text, pose, and image representations, ensuring consistent grounding across modalities. Furthermore, we design an iterative scene construction scheme, progressively generating complex multi-human interactions while effectively decomposing the overall generation complexity. Extensive experiments demonstrate that our method substantially improves prompt alignment and scene diversity in multi-person image generation.

📄 PDF Abstract BibTeX arXiv:2605.23178

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

Iterative Refinement Improves Compositional Image Generation

2026-01-21 · Shantanu Jaiswal, Mihir Prabhudesai, Nikash Bhardwaj, Zheyang Qin 외 arxiv

Text-to-image (T2I) models have achieved remarkable progress, yet they continue to struggle with complex prompts that require simultaneously handling multiple objects, relations, and attributes. Existing inference-time s…

Image Generation

Compositional generalization through meta sequence-to-sequence learning

2019-06-12 · NeurIPS 2019 12 · Brenden M. Lake

People can learn a new concept and use it compositionally, understanding how to "blicket twice" after learning how to "blicket." In contrast, powerful sequence-to-sequence (seq2seq) neural networks fail such tests of com…

Holistic, Instance-Level Human Parsing

2017-09-11 · Qizhu Li, Anurag Arnab, Philip H. S. Torr

Object parsing -- the task of decomposing an object into its semantic parts -- has traditionally been formulated as a category-level segmentation problem. Consequently, when there are multiple objects in an image, curren…

Human DetectionHuman ParsingMulti-Human ParsingObject+1

Composing Ensembles of Pre-trained Models via Iterative Consensus

2022-10-20 · Shuang Li, Yilun Du, Joshua B. Tenenbaum, Antonio Torralba 외

Large pre-trained models exhibit distinct and complementary capabilities dependent on the data they are trained on. Language models such as GPT-3 are capable of textual reasoning but cannot understand visual information,…

Arithmetic ReasoningImage GenerationMathMathematical Reasoning+2

CP-decomposition with Tensor Power Method for Convolutional Neural Networks Compression

2017-01-25 · Marcella Astrid, Seung-Ik Lee

Convolutional Neural Networks (CNNs) has shown a great success in many areas including complex image classification tasks. However, they need a lot of memory and computational cost, which hinders them from running in rel…

General Classificationimage-classificationImage Classification