paper-with-me

홈 › Papers

MagicScroll: Nontypical Aspect-Ratio Image Generation for Visual Storytelling via Multi-Layered Semantic-Aware Denoising

2023-12-18 · Bingyuan Wang, Hengyu Meng, Zeyu Cai, Lanjiong Li, Yue Ma, Qifeng Chen, Zeyu Wang

Visual storytelling often uses nontypical aspect-ratio images like scroll paintings, comic strips, and panoramas to create an expressive and compelling narrative. While generative AI has achieved great success and shown the potential to reshape the creative industry, it remains a challenge to generate coherent and engaging content with arbitrary size and controllable style, concept, and layout, all of which are essential for visual storytelling. To overcome the shortcomings of previous methods including repetitive content, style inconsistency, and lack of controllability, we propose MagicScroll, a multi-layered, progressive diffusion-based image generation framework with a novel semantic-aware denoising process. The model enables fine-grained control over the generated image on object, scene, and background levels with text, image, and layout conditions. We also establish the first benchmark for nontypical aspect-ratio image generation for visual storytelling including mediums like paintings, comics, and cinematic panoramas, with customized metrics for systematic evaluation. Through comparative and ablation studies, MagicScroll showcases promising results in aligning with the narrative text, improving visual coherence, and engaging the audience. We plan to release the code and benchmark in the hope of a better collaboration between AI researchers and creative practitioners involving visual storytelling.

📄 PDF Abstract BibTeX arXiv:2312.10899

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingImage GenerationVisual Storytelling

Similar Papers 제목 키워드 기반

Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

2024-02-27 · Daiqing Li, Aleks Kamko, Ehsan Akhgari, Ali Sabet 외

In this work, we share three insights for achieving state-of-the-art aesthetic quality in text-to-image generative models. We focus on three critical aspects for model improvement: enhancing color and contrast, improving…

Image GenerationText to Image GenerationText-to-Image Generation

Deciphering Personalization: Towards Fine-Grained Explainability in Natural Language for Personalized Image Generation Models

2025-11-02 · Haoming Wang, Wei Gao arxiv

Image generation models are usually personalized in practical uses in order to better meet the individual users' heterogeneous needs, but most personalized models lack explainability about how they are being personalized…

Personalized Image Generation

FRAbench and GenEval: Scaling Fine-Grained Aspect Evaluation across Tasks, Modalities

2025-05-19 · Shibo Hong, Jiahao Ying, Haiyuan Liang, Mengdi Zhang 외

Evaluating the open-ended outputs of large language models (LLMs) has become a bottleneck as model capabilities, task diversity, and modality coverage rapidly expand. Existing "LLM-as-a-Judge" evaluators are typically na…

Image GenerationText Generation

ElasticDiffusion: Training-free Arbitrary Size Image Generation through Global-Local Content Separation

2023-11-30 · CVPR 2024 1 · Moayed Haji-Ali, Guha Balakrishnan, Vicente Ordonez

Diffusion models have revolutionized image generation in recent years, yet they are still limited to a few sizes and aspect ratios. We propose ElasticDiffusion, a novel training-free decoding method that enables pretrain…

Image Generation

An Experience-based Direct Generation approach to Automatic Image Cropping

2022-12-30 · Casper Christensen, Aneesh Vartakavi

Automatic Image Cropping is a challenging task with many practical downstream applications. The task is often divided into sub-problems - generating cropping candidates, finding the visually important regions, and determ…

Image Cropping