paper-with-me

Papers

SpotActor: Training-Free Layout-Controlled Consistent Image Generation

2024-09-07 · Jiahao Wang, Caixia Yan, Weizhan Zhang, Haonan Lin, Mengmeng Wang, Guang Dai, Tieliang Gong, Hao Sun, Jingdong Wang

Text-to-image diffusion models significantly enhance the efficiency of artistic creation with high-fidelity image generation. However, in typical application scenarios like comic book production, they can neither place each subject into its expected spot nor maintain the consistent appearance of each subject across images. For these issues, we pioneer a novel task, Layout-to-Consistent-Image (L2CI) generation, which produces consistent and compositional images in accordance with the given layout conditions and text prompts. To accomplish this challenging task, we present a new formalization of dual energy guidance with optimization in a dual semantic-latent space and thus propose a training-free pipeline, SpotActor, which features a layout-conditioned backward update stage and a consistent forward sampling stage. In the backward stage, we innovate a nuanced layout energy function to mimic the attention activations with a sigmoid-like objective. While in the forward stage, we design Regional Interconnection Self-Attention (RISA) and Semantic Fusion Cross-Attention (SFCA) mechanisms that allow mutual interactions across images. To evaluate the performance, we present ActorBench, a specified benchmark with hundreds of reasonable prompt-box pairs stemming from object detection datasets. Comprehensive experiments are conducted to demonstrate the effectiveness of our method. The results prove that SpotActor fulfills the expectations of this task and showcases the potential for practical applications with superior layout alignment, subject consistency, prompt conformity and background diversity.

📄 PDF Abstract BibTeX arXiv:2409.04801

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generationobject-detectionObject Detection

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

FreeMorph: Tuning-Free Generalized Image Morphing with Diffusion Model

2025-07-02 · Yukang Cao, Chenyang Si, Jinghao Wang, Ziwei Liu

We present FreeMorph, the first tuning-free method for image morphing that accommodates inputs with different semantics or layouts. Unlike existing methods that rely on finetuning pre-trained diffusion models and are lim…

DenoisingImage Morphing

Zero-Painter: Training-Free Layout Control for Text-to-Image Synthesis

2024-06-06 · CVPR 2024 1 · Marianna Ohanyan, Hayk Manukyan, Zhangyang Wang, Shant Navasardyan 외

We present Zero-Painter, a novel training-free framework for layout-conditional text-to-image synthesis that facilitates the creation of detailed and controlled imagery from textual prompts. Our method utilizes object ma…

Conditional Text-to-Image SynthesisImage Generation

ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion Transformers

2025-12-01 · Yiyang Ma, Feng Zhou, Xuedan Yin, Pu Cao 외 arxiv

Leveraging pre-trained Diffusion Transformers (DiTs) for high-resolution (HR) image synthesis often leads to spatial layout collapse and degraded texture fidelity. Prior work mitigates these issues with complex pipelines…

ConsistCompose: Unified Multimodal Layout Control for Image Composition

2025-11-23 · Xuanke Shi, Boxuan Li, Xiaoyang Han, Zhongang Cai 외 arxiv

Unified multimodal models that couple visual understanding with image generation have advanced rapidly, yet most systems still focus on visual grounding-aligning language with image regions-while their generative counter…

Visual GroundingImage Generation

LAMIC: Layout-Aware Multi-Image Composition via Scalability of Multimodal Diffusion Transformer

2025-08-01 · Yuzhuo Chen, Zehua Ma, Jianhua Wang, Kai Kang 외 arxiv

In controllable image synthesis, generating coherent and consistent images from multiple references with spatial layout awareness remains an open challenge. We present LAMIC, a Layout-Aware Multi-Image Composition framew…

Zero-shot Generalization