paper-with-me

홈 › Papers

One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt

2025-01-23 · Tao Liu, Kai Wang, Senmao Li, Joost Van de Weijer, Fahad Shahbaz Khan, Shiqi Yang, Yaxing Wang, Jian Yang, Ming-Ming Cheng

Text-to-image generation models can create high-quality images from input prompts. However, they struggle to support the consistent generation of identity-preserving requirements for storytelling. Existing approaches to this problem typically require extensive training in large datasets or additional modifications to the original model architectures. This limits their applicability across different domains and diverse diffusion model configurations. In this paper, we first observe the inherent capability of language models, coined context consistency, to comprehend identity through context with a single prompt. Drawing inspiration from the inherent context consistency, we propose a novel training-free method for consistent text-to-image (T2I) generation, termed "One-Prompt-One-Story" (1Prompt1Story). Our approach 1Prompt1Story concatenates all prompts into a single input for T2I diffusion models, initially preserving character identities. We then refine the generation process using two novel techniques: Singular-Value Reweighting and Identity-Preserving Cross-Attention, ensuring better alignment with the input description for each frame. In our experiments, we compare our method against various existing consistent T2I generation approaches to demonstrate its effectiveness through quantitative metrics and qualitative assessments. Code is available at https://github.com/byliutao/1Prompt1Story.

📄 PDF Abstract BibTeX arXiv:2501.13554

Code (1)

byliutao/1prompt1story 공식 구현 pytorch

Tasks

Image GenerationStory GenerationText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs

2025-09-30 · Jia Jun Cheng Xian, Muchen Li, Haotian Yang, Xin Tao 외 arxiv

Recent advances in diffusion-based text-to-image (T2I) models have led to remarkable success in generating high-quality images from textual prompts. However, ensuring accurate alignment between the text and the generated…

Reinforcement Learning

Infinite-Story: A Training-Free Consistent Text-to-Image Generation

2025-11-17 · Jihun Park, Kyoungmin Lee, Jongmin Gim, Hyeonseo Jo 외 arxiv

We present Infinite-Story, a training-free framework for consistent text-to-image (T2I) generation tailored for multi-prompt storytelling scenarios. Built upon a scale-wise autoregressive model, our method addresses two …

Text-to-Image GenerationVisual Storytelling

FreeStory: Training-Free Character Consistency for Free-Form Visual Storytelling

2026-06-23 · Sibo Dong, Ismail Shaheen, Sarah Adel Bargal arxiv

Visual storytelling aims to generate image sequences that are both aligned with narrative prompts and consistent in character appearance across images. Recent training-free methods improve character consistency by reusin…

Visual Storytelling

Sidecar: Training-Free Semantic Reuse for Character-Consistent Free-form Visual Storytelling

2026-08-27 · Sibo Dong, Sarah Adel Bargal arxiv

Visual storytelling requires generating images that follow a narrative while preserving consistent character identities across frames. In free-form story generation, a character is fully described only when first introdu…

Visual StorytellingStory Generation

DeCorStory: Gram-Schmidt Prompt Embedding Decorrelation for Consistent Storytelling

2026-02-01 · Ayushman Sarkar, Zhenyu Yu, Mohd Yamani Idna Idris arxiv

Maintaining visual and semantic consistency across frames is a key challenge in text-to-image storytelling. Existing training-free methods, such as One-Prompt-One-Story, concatenate all prompts into a single sequence, wh…