RoViST: Learning Robust Metrics for Visual Storytelling
Visual storytelling (VST) is the task of generating a story paragraph that describes a given image sequence. Most existing storytelling approaches have evaluated their models using traditional natural language generation metrics like BLEU or CIDEr. However, such metrics based on $n$-gram matching tend to have poor correlation with human evaluation scores and do not explicitly consider other criteria necessary for storytelling such as sentence structure or topic coherence. Moreover, a single score is not enough to assess a story as it does not inform us about what specific errors were made by the model. In this paper, we propose 3 evaluation metrics sets that analyses which aspects we would look for in a good story: 1) visual grounding, 2) coherence, and 3) non-redundancy. We measure the reliability of our metric sets by analysing its correlation with human judgement scores on a sample of machine stories obtained from 4 state-of-the-arts models trained on the Visual Storytelling Dataset (VIST). Our metric sets outperforms other metrics on human correlation, and could be served as a learning based evaluation metric set that is complementary to existing rule-based metrics.
Code (0)
등록된 구현이 없습니다.
Tasks
SentenceText GenerationVisual GroundingVisual StorytellingSimilar Papers 제목 키워드 기반
VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
Visual storytelling is an interdisciplinary field combining computer vision and natural language processing to generate cohesive narratives from sequences of images. This paper presents a novel approach that leverages re…
Visual GroundingVisual StorytellingRoViST:Learning Robust Metrics for Visual Storytelling
Visual storytelling (VST) is the task of generating a story paragraph that describes a given image sequence. Most existing storytelling approaches have evaluated their models using traditional natural language generation…
SentenceText GenerationVisual GroundingVisual StorytellingRoViST: Learning Robust Metrics for Visual Storytelling
Visual storytelling (VST) is the task of generating a story paragraph that describes a given image sequence. Most existing storytelling approaches have evaluated their models using traditional natural language generation…
SentenceText GenerationVisual GroundingVisual StorytellingHide-and-Tell: Learning to Bridge Photo Streams for Visual Storytelling
Visual storytelling is a task of creating a short story based on photo streams. Unlike existing visual captioning, storytelling aims to contain not only factual descriptions, but also human-like narration and semantics. …
Image CaptioningVisual StorytellingA Hierarchical Approach for Visual Storytelling Using Image Description
One of the primary challenges of visual storytelling is developing techniques that can maintain the context of the story over long event sequences to generate human-like stories. In this paper, we propose a hierarchical …
DecoderImage DescriptionSentenceVisual Storytelling