paper-with-me

Papers

Learning to Model Multimodal Semantic Alignment for Story Visualization

2022-11-14 · Bowen Li, Thomas Lukasiewicz

Story visualization aims to generate a sequence of images to narrate each sentence in a multi-sentence story, where the images should be realistic and keep global consistency across dynamic scenes and characters. Current works face the problem of semantic misalignment because of their fixed architecture and diversity of input modalities. To address this problem, we explore the semantic alignment between text and image representations by learning to match their semantic levels in the GAN-based generative model. More specifically, we introduce dynamic interactions according to learning to dynamically explore various semantic depths and fuse the different-modal information at a matched semantic level, which thus relieves the text-image semantic misalignment problem. Extensive experiments on different datasets demonstrate the improvements of our approach, neither using segmentation masks nor auxiliary captioning networks, on image quality and story consistency, compared with state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2211.07289

Code (0)

등록된 구현이 없습니다.

Tasks

DiversitySentenceStory Visualization

Similar Papers 제목 키워드 기반

Generating a Temporally Coherent Image Sequence for a Story by Multimodal Recurrent Transformers

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Story visualization is a challenging text-to-image generation task for the difficulty of rendering visual details from abstract text descriptions. Besides the difficulty of image generation, the generator also need to c…

Image GenerationSentenceStory VisualizationText to Image Generation+1

DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion

2024-07-17 · Huiguo He, Huan Yang, Zixi Tuo, Yuan Zhou 외

Story visualization aims to create visually compelling images or videos corresponding to textual narratives. Despite recent advances in diffusion models yielding promising results, existing methods still struggle to crea…

DescriptiveStory Visualization

Dialogue Director: Bridging the Gap in Dialogue Visualization for Multimodal Storytelling

2024-12-30 · Min Zhang, Zilin Wang, Liyan Chen, KunHong Liu 외

Recent advances in AI-driven storytelling have enhanced video generation and story visualization. However, translating dialogue-centric scripts into coherent storyboards remains a significant challenge due to limited scr…

Retrieval-augmented GenerationStory VisualizationVideo Generation

StoryWeaver: A Unified World Model for Knowledge-Enhanced Story Character Customization

2024-12-10 · Jinlu Zhang, Jiji Tang, Rongsheng Zhang, Tangjie Lv 외

Story visualization has gained increasing attention in artificial intelligence. However, existing methods still struggle with maintaining a balance between character identity preservation and text-semantics alignment, la…

Story Visualization

Generating a Temporally Coherent Visual Story by Multimodal Recurrent Transformers

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Story visualization is a challenging text-to-image generation task for the difficulty of rendering visual details from abstract text descriptions. Besides the difficulty of image generation, the generator also needs to c…

Image GenerationSentenceStory VisualizationText to Image Generation+1