Learning to Model Multimodal Semantic Alignment for Story Visualization
Story visualization aims to generate a sequence of images to narrate each sentence in a multi-sentence story, where the images should be realistic and keep global consistency across dynamic scenes and characters. Current works face the problem of semantic misalignment because of their fixed architecture and diversity of input modalities. To address this problem, we explore the semantic alignment between text and image representations by learning to match their semantic levels in the GAN-based generative model. More specifically, we introduce dynamic interactions according to learning to dynamically explore various semantic depths and fuse the different-modal information at a matched semantic level, which thus relieves the text-image semantic misalignment problem. Extensive experiments on different datasets demonstrate the improvements of our approach, neither using segmentation masks nor auxiliary captioning networks, on image quality and story consistency, compared with state-of-the-art methods.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversitySentenceStory VisualizationSimilar Papers 제목 키워드 기반
Generating a Temporally Coherent Image Sequence for a Story by Multimodal Recurrent Transformers
Story visualization is a challenging text-to-image generation task for the difficulty of rendering visual details from abstract text descriptions. Besides the difficulty of image generation, the generator also need to c…
Image GenerationSentenceStory VisualizationText to Image Generation+1DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion
Story visualization aims to create visually compelling images or videos corresponding to textual narratives. Despite recent advances in diffusion models yielding promising results, existing methods still struggle to crea…
DescriptiveStory VisualizationDialogue Director: Bridging the Gap in Dialogue Visualization for Multimodal Storytelling
Recent advances in AI-driven storytelling have enhanced video generation and story visualization. However, translating dialogue-centric scripts into coherent storyboards remains a significant challenge due to limited scr…
Retrieval-augmented GenerationStory VisualizationVideo GenerationStoryWeaver: A Unified World Model for Knowledge-Enhanced Story Character Customization
Story visualization has gained increasing attention in artificial intelligence. However, existing methods still struggle with maintaining a balance between character identity preservation and text-semantics alignment, la…
Story VisualizationGenerating a Temporally Coherent Visual Story by Multimodal Recurrent Transformers
Story visualization is a challenging text-to-image generation task for the difficulty of rendering visual details from abstract text descriptions. Besides the difficulty of image generation, the generator also needs to c…
Image GenerationSentenceStory VisualizationText to Image Generation+1