Generating a Temporally Coherent Image Sequence for a Story by Multimodal Recurrent Transformers
Story visualization is a challenging text-to-image generation task for the difficulty of rendering visual details from abstract text descriptions. Besides the difficulty of image generation, the generator also need to conform to the narrative of a multi-sentence story input. While prior arts in this domain has focused on improving semantic relevance between generated images and input text, controlling the generated images to be temporally consistent still remains as a challenge. To generate a semantically coherent image sequence, we propose an explicit memory controller which can augment the temporal coherence of images in the multi-modal autoregressive transformer, and call Story visualization by MultimodAl Recurrent Transformers or SMART for short. Our method generates high resolution high quality images, outperforming prior works by a significant margin across multiple evaluation metrics on PororoSV dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
Image GenerationSentenceStory VisualizationText to Image GenerationText-to-Image GenerationSimilar Papers 제목 키워드 기반
Generating a Temporally Coherent Visual Story by Multimodal Recurrent Transformers
Story visualization is a challenging text-to-image generation task for the difficulty of rendering visual details from abstract text descriptions. Besides the difficulty of image generation, the generator also needs to c…
Image GenerationSentenceStory VisualizationText to Image Generation+1Intelligent Grimm -- Open-ended Visual Storytelling via Latent Diffusion Models
Generative models have recently exhibited exceptional capabilities in text-to-image generation, but still struggle to generate image sequences coherently. In this work, we focus on a novel, yet challenging task of genera…
Image GenerationStory VisualizationStyle TransferText to Image Generation+2Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion Models
Generative models have recently exhibited exceptional capabilities in text-to-image generation but still struggle to generate image sequences coherently. In this work we focus on a novel yet challenging task of gener…
Image GenerationText to Image GenerationText-to-Image GenerationVisual StorytellingCharacter-Preserving Coherent Story Visualization
Story visualization aims at generating a sequence of images to narrate each sentence in a multi-sentence story. Different from video generation that focuses on maintaining the continuity of generated images (frames), sto…
Representation LearningSentenceStory VisualizationHierarchically Structured Reinforcement Learning for Topically Coherent Visual Story Generation
We propose a hierarchically structured reinforcement learning approach to address the challenges of planning for generating coherent multi-sentence stories for the visual storytelling task. Within our framework, the task…
DecoderDeep Reinforcement Learningreinforcement-learningReinforcement Learning+4