paper-with-me

홈 › Papers

Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning

2024-08-12 · Yingjin Song, Denis Paperno, Albert Gatt

Visual storytelling systems generate multi-sentence stories from image sequences. In this task, capturing contextual information and bridging visual variation bring additional challenges. We propose a simple yet effective framework that leverages the generalization capabilities of pretrained foundation models, only training a lightweight vision-language mapping network to connect modalities, while incorporating context to enhance coherence. We introduce a multimodal contrastive objective that also improves visual relevance and story informativeness. Extensive experimental results, across both automatic metrics and human evaluations, demonstrate that the stories generated by our framework are diverse, coherent, informative, and interesting.

📄 PDF Abstract BibTeX arXiv:2408.06259

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningInformativenessSentenceVisual Storytelling

Similar Papers 제목 키워드 기반

Visual Storytelling with Question-Answer Plans

2023-10-08 · Danyang Liu, Mirella Lapata, Frank Keller

Visual storytelling aims to generate compelling narratives from image sequences. Existing models often focus on enhancing the representation of the image sequence, e.g., with external knowledge sources or advanced graph …

Visual Storytelling

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?

2025-04-27 · Mohamed Gado, Towhid Taliee, Muhammad Memon, Dmitry Ignatov 외

Visual storytelling is an interdisciplinary field combining computer vision and natural language processing to generate cohesive narratives from sequences of images. This paper presents a novel approach that leverages re…

Visual GroundingVisual Storytelling

TARN-VIST: Topic Aware Reinforcement Network for Visual Storytelling

2024-03-18 · Weiran Chen, Xin Li, Jiaqi Su, Guiqian Zhu 외

As a cross-modal task, visual storytelling aims to generate a story for an ordered image sequence automatically. Different from the image captioning task, visual storytelling requires not only modeling the relationships …

Image CaptioningVisual Storytelling

MagicScroll: Nontypical Aspect-Ratio Image Generation for Visual Storytelling via Multi-Layered Semantic-Aware Denoising

2023-12-18 · Bingyuan Wang, Hengyu Meng, Zeyu Cai, Lanjiong Li 외

Visual storytelling often uses nontypical aspect-ratio images like scroll paintings, comic strips, and panoramas to create an expressive and compelling narrative. While generative AI has achieved great success and shown …

DenoisingImage GenerationVisual Storytelling

A Hierarchical Approach for Visual Storytelling Using Image Description

2019-09-26 · Md Sultan Al Nahian, Tasmia Tasrin, Sagar Gandhi, Ryan Gaines 외

One of the primary challenges of visual storytelling is developing techniques that can maintain the context of the story over long event sequences to generate human-like stories. In this paper, we propose a hierarchical …

DecoderImage DescriptionSentenceVisual Storytelling