paper-with-me

홈 › Papers

Multimodal Event Transformer for Image-guided Story Ending Generation

2023-01-26 · Yucheng Zhou, Guodong Long

Image-guided story ending generation (IgSEG) is to generate a story ending based on given story plots and ending image. Existing methods focus on cross-modal feature fusion but overlook reasoning and mining implicit information from story plots and ending image. To tackle this drawback, we propose a multimodal event transformer, an event-based reasoning framework for IgSEG. Specifically, we construct visual and semantic event graphs from story plots and ending image, and leverage event-based reasoning to reason and mine implicit information in a single modality. Next, we connect visual and semantic event graphs and utilize cross-modal fusion to integrate different-modality features. In addition, we propose a multimodal injector to adaptive pass essential information to decoder. Besides, we present an incoherence detection to enhance the understanding context of a story plot and the robustness of graph modeling for our model. Experimental results show that our method achieves state-of-the-art performance for the image-guided story ending generation.

📄 PDF Abstract BibTeX arXiv:2301.11357

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderImage-guided Story Ending Generation

Similar Papers 제목 키워드 기반

MMT: Image-guided Story Ending Generation with Multimodal Memory Transformer

2022-10-10 · ACM MM 2022 10 · Dizhan Xue, Shengsheng Qian, Quan Fang, Changsheng Xu

As a specific form of story generation, Image-guided Story Ending Generation (IgSEG) is a recently proposed task of generating a story ending for a given multi-sentence story plot and an ending-related image. Unlike exis…

DecoderImage CaptioningImage-guided Story Ending GenerationSentence+1

Iterative Adversarial Attack on Image-guided Story Ending Generation

2023-05-16 · Youze Wang, WenBo Hu, Richang Hong

Multimodal learning involves developing models that can integrate information from various sources like images and texts. In this field, multimodal text generation is a crucial aspect that involves processing data from m…

Adversarial AttackAdversarial RobustnessAdversarial TextImage-guided Story Ending Generation+4

Generating a Temporally Coherent Image Sequence for a Story by Multimodal Recurrent Transformers

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Story visualization is a challenging text-to-image generation task for the difficulty of rendering visual details from abstract text descriptions. Besides the difficulty of image generation, the generator also need to c…

Image GenerationSentenceStory VisualizationText to Image Generation+1

DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion

2024-07-17 · Huiguo He, Huan Yang, Zixi Tuo, Yuan Zhou 외

Story visualization aims to create visually compelling images or videos corresponding to textual narratives. Despite recent advances in diffusion models yielding promising results, existing methods still struggle to crea…

DescriptiveStory Visualization

Multimodal Incremental Transformer with Visual Grounding for Visual Dialogue Generation

2021-09-17 · Findings (ACL) 2021 8 · Feilong Chen, Fandong Meng, Xiuyi Chen, Peng Li 외

Visual dialogue is a challenging task since it needs to answer a series of coherent questions on the basis of understanding the visual environment. Previous studies focus on the implicit exploration of multimodal co-refe…

Dialogue GenerationVisual Grounding