paper-with-me

홈 › Papers

Generating a Temporally Coherent Image Sequence for a Story by Multimodal Recurrent Transformers

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Story visualization is a challenging text-to-image generation task for the difficulty of rendering visual details from abstract text descriptions. Besides the difficulty of image generation, the generator also need to conform to the narrative of a multi-sentence story input. While prior arts in this domain has focused on improving semantic relevance between generated images and input text, controlling the generated images to be temporally consistent still remains as a challenge. To generate a semantically coherent image sequence, we propose an explicit memory controller which can augment the temporal coherence of images in the multi-modal autoregressive transformer, and call Story visualization by MultimodAl Recurrent Transformers or SMART for short. Our method generates high resolution high quality images, outperforming prior works by a significant margin across multiple evaluation metrics on PororoSV dataset.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationSentenceStory VisualizationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

Generating a Temporally Coherent Visual Story by Multimodal Recurrent Transformers

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Story visualization is a challenging text-to-image generation task for the difficulty of rendering visual details from abstract text descriptions. Besides the difficulty of image generation, the generator also needs to c…

Image GenerationSentenceStory VisualizationText to Image Generation+1

Intelligent Grimm -- Open-ended Visual Storytelling via Latent Diffusion Models

2023-06-01 · Chang Liu, HaoNing Wu, Yujie Zhong, Xiaoyun Zhang 외

Generative models have recently exhibited exceptional capabilities in text-to-image generation, but still struggle to generate image sequences coherently. In this work, we focus on a novel, yet challenging task of genera…

Image GenerationStory VisualizationStyle TransferText to Image Generation+2

Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion Models

2024-01-01 · CVPR 2024 1 · Chang Liu, HaoNing Wu, Yujie Zhong, Xiaoyun Zhang 외

Generative models have recently exhibited exceptional capabilities in text-to-image generation but still struggle to generate image sequences coherently. In this work we focus on a novel yet challenging task of gener…

Image GenerationText to Image GenerationText-to-Image GenerationVisual Storytelling

Character-Preserving Coherent Story Visualization

2020-08-01 · ECCV 2020 8 · Yun-Zhu Song, Zhi Rui Tam, Hung-Jen Chen, Huiao-Han Lu 외

Story visualization aims at generating a sequence of images to narrate each sentence in a multi-sentence story. Different from video generation that focuses on maintaining the continuity of generated images (frames), sto…

Representation LearningSentenceStory Visualization

Hierarchically Structured Reinforcement Learning for Topically Coherent Visual Story Generation

2018-05-21 · Qiuyuan Huang, Zhe Gan, Asli Celikyilmaz, Dapeng Wu 외

We propose a hierarchically structured reinforcement learning approach to address the challenges of planning for generating coherent multi-sentence stories for the visual storytelling task. Within our framework, the task…

DecoderDeep Reinforcement Learningreinforcement-learningReinforcement Learning+4