paper-with-me

홈 › Papers

Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation

2025-08-12 · Ao Ma, Jiasong Feng, Ke Cao, Jing Wang, Yun Wang, Quanwei Zhang, Zhanjie Zhang arxiv

Storytelling tasks involving generating consistent subjects have gained significant attention recently. However, existing methods, whether training-free or training-based, continue to face challenges in maintaining subject consistency due to the lack of fine-grained guidance and inter-frame interaction. Additionally, the scarcity of high-quality data in this field makes it difficult to precisely control storytelling tasks, including the subject's position, appearance, clothing, expression, and posture, thereby hindering further advancements. In this paper, we demonstrate that layout conditions, such as the subject's position and detailed attributes, effectively facilitate fine-grained interactions between frames. This not only strengthens the consistency of the generated frame sequence but also allows for precise control over the subject's position, appearance, and other key details. Building on this, we introduce an advanced storytelling task: Layout-Togglable Storytelling, which enables precise subject control by incorporating layout conditions. To address the lack of high-quality datasets with layout annotations for this task, we develop Lay2Story-1M, which contains over 1 million 720p and higher-resolution images, processed from approximately 11,300 hours of cartoon videos. Building on Lay2Story-1M, we create Lay2Story-Bench, a benchmark with 3,000 prompts designed to evaluate the performance of different methods on this task. Furthermore, we propose Lay2Story, a robust framework based on the Diffusion Transformers (DiTs) architecture for Layout-Togglable Storytelling tasks. Through both qualitative and quantitative experiments, we find that our method outperforms the previous state-of-the-art (SOTA) techniques, achieving the best results in terms of consistency, semantic correlation, and aesthetic quality.

📄 PDF Abstract BibTeX arXiv:2508.08949

Code (0)

등록된 구현이 없습니다.

Tasks

Story Generation

Similar Papers 제목 키워드 기반

Dolfin: Diffusion Layout Transformers without Autoencoder

2023-10-25 · Yilin Wang, Zeyuan Chen, Liangjun Zhong, Zheng Ding 외

In this paper, we introduce a novel generative model, Diffusion Layout Transformers without Autoencoder (Dolfin), which significantly improves the modeling capability with reduced complexity compared to existing methods.…

Layout Generation

DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video Models

2025-12-01 · Patrick Kwon, Chen Chen arxiv

Current story visualization methods tend to position subjects solely by text and face challenges in maintaining artistic consistency. To address these limitations, we introduce DreamingComics, a layout-aware story visual…

Story Visualization

Manga Generation via Layout-controllable Diffusion

2024-12-26 · Siyu Chen, Dengjie Li, Zenghao Bao, Yao Zhou 외

Generating comics through text is widely studied. However, there are few studies on generating multi-panel Manga (Japanese comics) solely based on plain text. Japanese manga contains multiple panels on a single page, wit…

Semantic correspondence

LEARN: A Story-Driven Layout-to-Image Generation Framework for STEM Instruction

2025-08-15 · Maoquan Zhang, Bisser Raytchev, Xiujuan Sun arxiv

LEARN is a layout-aware diffusion framework designed to generate pedagogically aligned illustrations for STEM education. It leverages a curated BookCover dataset that provides narrative layouts and structured visual cues…

Layout-to-Image GenerationKnowledge Graphs

CogCartoon: Towards Practical Story Visualization

2023-12-17 · Zhongyang Zhu, Jie Tang

The state-of-the-art methods for story visualization demonstrate a significant demand for training data and storage, as well as limited flexibility in story presentation, thereby rendering them impractical for real-world…

Story Visualization