paper-with-me

홈 › Papers

S2ED: From Story to Executable Descriptions for Consistency-Aware Story Illustration

2026-05-21 · Sijing Yin, Jiamou Liu, Xiao Tang, Yaser Shakib, Qian Liu arxiv

Multi-frame story illustration requires long-horizon coherence beyond single-image text-to-image generation, including narrative decomposition and persistent character identity, layout, and affect across frames. We propose Story-to-Executable Descriptions (S2ED), a training-free, model-agnostic, prompt-layer framework that converts a full story into a sequence of explicit, editable executable descriptions for more consistent rendering. S2ED coordinates three agents to segment the narrative, ground canonical character attributes, and enrich spatial and affective cues, enabling interpretable prompt-carried state propagation and local edits to repair drift without retraining the generator. Experiments on Flintstones and Shakoo Maku show that S2ED improves sequence-level consistency and character fidelity over strong prompting, large-model planning, and a reference training-based method, under both automatic metrics and human judgments. We also deploy S2ED in an end-to-end story-to-storybook system for children's illustrated stories, with a supplementary video.

📄 PDF Abstract BibTeX arXiv:2605.22448

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image Generation

Similar Papers 제목 키워드 기반

FreeStory: Training-Free Character Consistency for Free-Form Visual Storytelling

2026-06-23 · Sibo Dong, Ismail Shaheen, Sarah Adel Bargal arxiv

Visual storytelling aims to generate image sequences that are both aligned with narrative prompts and consistent in character appearance across images. Recent training-free methods improve character consistency by reusin…

Visual Storytelling

Customized Visual Storytelling with Unified Multimodal LLMs

2026-03-29 · Wei-Hua Li, Cheng Sun, Chu-Song Chen arxiv

Multimodal story customization aims to generate coherent story flows conditioned on textual descriptions, reference identity images, and shot types. While recent progress in story generation has shown promising results, …

Visual StorytellingStory Generation

TaleDiffusion: Multi-Character Story Generation with Dialogue Rendering

2025-09-04 · Ayan Banerjee, Josep Llados, Umapada Pal, Anjan Dutta arxiv

Text-to-story visualization is challenging due to the need for consistent interaction among multiple characters across frames. Existing methods struggle with character consistency, leading to artifact generation and inac…

Story VisualizationStory Generation

StoryTeller: Improving Long Video Description through Global Audio-Visual Character Identification

2024-11-11 · Yichen He, Yuan Lin, Jianchao Wu, Hanchong Zhang 외

Existing large vision-language models (LVLMs) are largely limited to processing short, seconds-long videos and struggle with generating coherent descriptions for extended video spanning minutes or more. Long video descri…

Large Language ModelMultimodal Large Language ModelMultiple-choiceVideo Description

ConWriter: Transition-Constrained Stateful Long-Form Story Generation with Lightweight Neuro-Symbolic Consistency Control

2026-05-27 · Jindong Li, Yang Yang, Zihao Liu, Yutao Yue 외 arxiv

Long-form story generation requires models to preserve narrative consistency across extended contexts, yet existing prompting-based methods often accumulate temporal, factual, character, commonsense, and stylistic errors…