paper-with-me

홈 › Papers

GROOViST: A Metric for Grounding Objects in Visual Storytelling

2023-10-26 · Aditya K Surikuchi, Sandro Pezzelle, Raquel Fernández

A proper evaluation of stories generated for a sequence of images -- the task commonly referred to as visual storytelling -- must consider multiple aspects, such as coherence, grammatical correctness, and visual grounding. In this work, we focus on evaluating the degree of grounding, that is, the extent to which a story is about the entities shown in the images. We analyze current metrics, both designed for this purpose and for general vision-text alignment. Given their observed shortcomings, we propose a novel evaluation tool, GROOViST, that accounts for cross-modal dependencies, temporal misalignments (the fact that the order in which entities appear in the story and the image sequence may not match), and human intuitions on visual grounding. An additional advantage of GROOViST is its modular design, where the contribution of each component can be assessed and interpreted individually.

📄 PDF Abstract BibTeX arXiv:2310.17770

Code (1)

akskuchi/groovist 공식 구현 pytorch

Tasks

Visual GroundingVisual Storytelling

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?

2025-04-27 · Mohamed Gado, Towhid Taliee, Muhammad Memon, Dmitry Ignatov 외

Visual storytelling is an interdisciplinary field combining computer vision and natural language processing to generate cohesive narratives from sequences of images. This paper presents a novel approach that leverages re…

Visual GroundingVisual Storytelling

Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition

2024-07-05 · Aditya K Surikuchi, Raquel Fernández, Sandro Pezzelle

Visual storytelling consists in generating a natural language story given a temporally ordered sequence of images. This task is not only challenging for models, but also very difficult to evaluate with automatic metrics …

Visual GroundingVisual Storytelling

RoViST:Learning Robust Metrics for Visual Storytelling

2022-05-08 · Eileen Wang, Caren Han, Josiah Poon

Visual storytelling (VST) is the task of generating a story paragraph that describes a given image sequence. Most existing storytelling approaches have evaluated their models using traditional natural language generation…

SentenceText GenerationVisual GroundingVisual Storytelling

RoViST: Learning Robust Metrics for Visual Storytelling

2022-07-01 · Findings (NAACL) 2022 7 · Eileen Wang, Caren Han, Josiah Poon

Visual storytelling (VST) is the task of generating a story paragraph that describes a given image sequence. Most existing storytelling approaches have evaluated their models using traditional natural language generation…

SentenceText GenerationVisual GroundingVisual Storytelling

RoViST: Learning Robust Metrics for Visual Storytelling

2021-12-17 · ACL ARR December 2022 12 · Anonymous

Visual storytelling (VST) is the task of generating a story paragraph that describes a given image sequence. Most existing storytelling approaches have evaluated their models using traditional natural language generation…

SentenceText GenerationVisual GroundingVisual Storytelling