paper-with-me

홈 › Papers

Modular StoryGAN with Background and Theme Awareness for Story Visualization

2022-06-02 · ICPRAI 2022 (3rd International Conference on Pattern Recognition and Artificial Intelligence) 2022 6 · Gábor Szűcs, Modafar Al-Shouha

Story visualization is a novel topic that intersects computer vision and natural language processing. In this task, given a series of natural language sentences that compose a story, a sequence of images should be generated that correspond to the sentences. Prior works have introduced recurrent generative models which outperform text-to-image models on this task; however, local and global consistency is a challenging attribute of these solutions. For the improvement, we proposed a new modular model architecture named Modular StoryGAN containing the best promising components of prior works. To measure the local and global consistency we introduced background and theme awareness, which are expected attributes of the solutions. Based on the human evaluation, the generated images demonstrate that Modular StoryGAN possesses background and theme awareness. Besides the subjective evaluation, the objective one also shows that our model outperforms the state-of-the-art CP-CSV and DuCo models.

📄 PDF Abstract BibTeX

Code (1)

modafarshouha/ModularStoryGAN 공식 구현 pytorch

Tasks

AttributeImage GenerationSegmentationStory VisualizationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Dogecoin Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

StoryGAN: A Sequential Conditional GAN for Story Visualization

2018-12-06 · CVPR 2019 6 · Yitong Li, Zhe Gan, Yelong Shen, Jingjing Liu 외

We propose a new task, called Story Visualization. Given a multi-sentence paragraph, the story is visualized by generating a sequence of images, one for each sentence. In contrast to video generation, story visualization…

SentenceStory VisualizationVideo Generation

Integrating Visuospatial, Linguistic and Commonsense Structure into Story Visualization

2021-10-21 · Adyasha Maharana, Mohit Bansal

While much research has been done in text-to-image synthesis, little work has been done to explore the usage of linguistic structure of the input text. Such information is even more important for story visualization sinc…

Dense CaptioningImage GenerationStory Visualization

MoPS: Modular Story Premise Synthesis for Open-Ended Automatic Story Generation

2024-06-09 · Yan Ma, Yu Qiao, PengFei Liu

A story premise succinctly defines a story's main idea, foundation, and trajectory. It serves as the initial trigger in automatic story generation. Existing sources of story premises are limited by a lack of diversity, u…

DiversitySentenceStory Generation

AWARE, Beyond Sentence Boundaries: A Contextual Transformer Framework for Identifying Cultural Capital in STEM Narratives

2025-10-06 · Khalid Mehtab Khan, Anagha Kulkarni arxiv

Identifying cultural capital (CC) themes in student reflections can offer valuable insights that help foster equitable learning environments in classrooms. However, themes such as aspirational goals or family support are…

Text Classification

StoryDALL-E: Adapting Pretrained Text-to-Image Transformers for Story Continuation

2022-09-13 · Adyasha Maharana, Darryl Hannan, Mohit Bansal

Recent advances in text-to-image synthesis have led to large pretrained transformers with excellent capabilities to generate visualizations from a given text. However, these models are ill-suited for specialized tasks li…

Image GenerationStory ContinuationStory VisualizationVideo Captioning