paper-with-me

홈 › Papers

From Image Captioning to Visual Storytelling

2025-07-31 · Admitos Passadakis, Yingjin Song, Albert Gatt arxiv

Visual Storytelling is a challenging multimodal task between Vision & Language, where the purpose is to generate a story for a stream of images. Its difficulty lies on the fact that the story should be both grounded to the image sequence but also narrative and coherent. The aim of this work is to balance between these aspects, by treating Visual Storytelling as a superset of Image Captioning, an approach quite different compared to most of prior relevant studies. This means that we firstly employ a vision-to-language model for obtaining captions of the input images, and then, these captions are transformed into coherent narratives using language-to-language methods. Our multifarious evaluation shows that integrating captioning and storytelling under a unified framework, has a positive impact on the quality of the produced stories. In addition, compared to numerous previous studies, this approach accelerates training time and makes our framework readily reusable and reproducible by anyone interested. Lastly, we propose a new metric/tool, named ideality, that can be used to simulate how far some results are from an oracle model, and we apply it to emulate human-likeness in visual storytelling.

📄 PDF Abstract BibTeX arXiv:2508.14045

Code (0)

등록된 구현이 없습니다.

Tasks

Visual StorytellingImage Captioning

Similar Papers 제목 키워드 기반

Semantic Alignment for Multimodal Large Language Models

2024-08-23 · Tao Wu, Mengze Li, Jingyuan Chen, Wei Ji 외

Research on Multi-modal Large Language Models (MLLMs) towards the multi-image cross-modal instruction has received increasing attention and made significant progress, particularly in scenarios involving closely resemblin…

Large Language ModelVisual Storytelling

TARN-VIST: Topic Aware Reinforcement Network for Visual Storytelling

2024-03-18 · Weiran Chen, Xin Li, Jiaqi Su, Guiqian Zhu 외

As a cross-modal task, visual storytelling aims to generate a story for an ordered image sequence automatically. Different from the image captioning task, visual storytelling requires not only modeling the relationships …

Image CaptioningVisual Storytelling

A-CAP: Anticipation Captioning with Commonsense Knowledge

2023-04-13 · CVPR 2023 1 · Duc Minh Vo, Quoc-An Luong, Akihiro Sugimoto, Hideki Nakayama

Humans possess the capacity to reason about the future based on a sparse collection of visual cues acquired over time. In order to emulate this ability, we introduce a novel task called Anticipation Captioning, which gen…

Image CaptioningLanguage ModelingLanguage ModellingVisual Storytelling

Hide-and-Tell: Learning to Bridge Photo Streams for Visual Storytelling

2020-02-03 · Yunjae Jung, Dahun Kim, Sanghyun Woo, Kyung-Su Kim 외

Visual storytelling is a task of creating a short story based on photo streams. Unlike existing visual captioning, storytelling aims to contain not only factual descriptions, but also human-like narration and semantics. …

Image CaptioningVisual Storytelling

MMCOMET: A Large-Scale Multimodal Commonsense Knowledge Graph for Contextual Reasoning

2026-03-01 · Eileen Wang, Hiba Arnaout, Dhita Pratama, Shuo Yang 외 arxiv

We present MMCOMET, the first multimodal commonsense knowledge graph (MMKG) that integrates physical, social, and eventive knowledge. MMCOMET extends the ATOMIC2020 knowledge graph to include a visual dimension, through …

Visual StorytellingImage CaptioningImage Retrieval