paper-with-me

홈 › Papers

Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion Models

2024-01-01 · CVPR 2024 1 · Chang Liu, HaoNing Wu, Yujie Zhong, Xiaoyun Zhang, Yanfeng Wang, Weidi Xie

Generative models have recently exhibited exceptional capabilities in text-to-image generation but still struggle to generate image sequences coherently. In this work we focus on a novel yet challenging task of generating a coherent image sequence based on a given storyline denoted as open-ended visual storytelling. We make the following three contributions: (i) to fulfill the task of visual storytelling we propose a learning-based auto-regressive image generation model termed as StoryGen with a novel vision-language context module that enables to generate the current frame by conditioning on the corresponding text prompt and preceding image-caption pairs; (ii) to address the data shortage of visual storytelling we collect paired image-text sequences by sourcing from online videos and open-source E-books establishing processing pipeline for constructing a large-scale dataset with diverse characters storylines and artistic styles named StorySalon; (iii) Quantitative experiments and human evaluations have validated the superiority of our StoryGen where we show it can generalize to unseen characters without any optimization and generate image sequences with coherent content and consistent character. Code dataset and models are available at https://haoningwu3639.github.io/StoryGen_Webpage/

📄 PDF Abstract BibTeX

Code (1)

haoningwu3639/StoryGen 공식 구현 pytorch

Tasks

Image GenerationText to Image GenerationText-to-Image GenerationVisual Storytelling

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Intelligent Grimm -- Open-ended Visual Storytelling via Latent Diffusion Models

2023-06-01 · Chang Liu, HaoNing Wu, Yujie Zhong, Xiaoyun Zhang 외

Generative models have recently exhibited exceptional capabilities in text-to-image generation, but still struggle to generate image sequences coherently. In this work, we focus on a novel, yet challenging task of genera…

Image GenerationStory VisualizationStyle TransferText to Image Generation+2

AgriDoctor: A Multimodal Intelligent Assistant for Agriculture

2025-09-21 · Mingqing Zhang, Zhuoning Xu, Peijie Wang, Rongji Li 외 arxiv

Accurate crop disease diagnosis is essential for sustainable agriculture and global food security. Existing methods, which primarily rely on unimodal models such as image-based classifiers and object detectors, are limit…

Multimodal ReasoningDomain Adaptation

FairyTailor: A Multimodal Generative Framework for Storytelling

2021-07-13 · Eden Bensaid, Mauro Martino, Benjamin Hoover, Jacob Andreas 외

Storytelling is an open-ended task that entails creative thinking and requires a constant flow of ideas. Natural language generation (NLG) for storytelling is especially challenging because it requires the generated text…

Story GenerationText Generation

Incorporating Textual Evidence in Visual Storytelling

2019-11-21 · WS 2019 11 · Tianyi Li, Sujian Li

Previous work on visual storytelling mainly focused on exploring image sequence as evidence for storytelling and neglected textual evidence for guiding story generation. Motivated by human storytelling process which reca…

Object RecognitionStory GenerationVisual Storytelling

Grimm: A Plug-and-Play Perturbation Rectifier for Graph Neural Networks Defending against Poisoning Attacks

2024-12-11 · Ao Liu, Wenshan Li, Beibei Li, Wengang Ma 외

Recent studies have revealed the vulnerability of graph neural networks (GNNs) to adversarial poisoning attacks on node classification tasks. Current defensive methods require substituting the original GNNs with defense …

Adversarial RobustnessClassificationNode Classification