paper-with-me

Papers

What's in a Caption? Dataset-Specific Linguistic Diversity and Its Effect on Visual Description Models and Metrics

2022-05-12 · David M. Chan, Austin Myers, Sudheendra Vijayanarasimhan, David A. Ross, Bryan Seybold, John F. Canny

While there have been significant gains in the field of automated video description, the generalization performance of automated description models to novel domains remains a major barrier to using these systems in the real world. Most visual description methods are known to capture and exploit patterns in the training data leading to evaluation metric increases, but what are those patterns? In this work, we examine several popular visual description datasets, and capture, analyze, and understand the dataset-specific linguistic patterns that models exploit but do not generalize to new domains. At the token level, sample level, and dataset level, we find that caption diversity is a major driving factor behind the generation of generic and uninformative captions. We further show that state-of-the-art models even outperform held-out ground truth captions on modern metrics, and that this effect is an artifact of linguistic diversity in datasets. Understanding this linguistic diversity is key to building strong captioning models, we recommend several methods and approaches for maintaining diversity in the collection of new data, and dealing with the consequences of limited diversity when using current models and metrics.

📄 PDF Abstract BibTeX arXiv:2205.06253

Code (1)

cannylab/vdtk 공식 구현

Tasks

DiversityVideo Description

Similar Papers 제목 키워드 기반

Surprisal reveals diversity gaps in image captioning and different scorers change the story

2025-11-06 · Nikolai Ilinykh, Simon Dobnik arxiv

We quantify linguistic diversity in image captioning with surprisal variance - the spread of token-level negative log-probabilities within a caption set. On the MSCOCO test set, we compare five state-of-the-art vision-an…

Image Captioning

ADS-Cap: A Framework for Accurate and Diverse Stylized Captioning with Unpaired Stylistic Corpora

2023-08-02 · Kanzhi Cheng, Zheng Ma, Shi Zong, Jianbing Zhang 외

Generating visually grounded image captions with specific linguistic styles using unpaired stylistic corpora is a challenging task, especially since we expect stylized captions with a wide variety of stylistic patterns. …

Contrastive LearningDiversityImage Captioning

Deep soccer captioning with transformer: dataset, semantics-related losses, and multi-level evaluation

2022-02-11 · Ahmad Hammoudeh, Bastien Vanderplaetse, Stéphane Dupont

This work aims at generating captions for soccer videos using deep learning. In this context, this paper introduces a dataset, model, and triple-level evaluation. The dataset consists of 22k caption-clip pairs and three …

DiversityOptical Flow Estimation

Diversity as a By-Product: Goal-oriented Language Generation Leads to Linguistic Variation

2021-07-01 · SIGDIAL (ACL) 2021 7 · Simeon Schüz, Ting Han, Sina Zarrieß

The ability for variation in language use is necessary for speakers to achieve their conversational goals, for instance when referring to objects in visual environments. We argue that diversity should not be modelled as …

DiversityImage CaptioningText Generation

VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph Captioning

2022-11-28 · Kashu Yamazaki, Khoa Vo, Sang Truong, Bhiksha Raj 외

Video paragraph captioning aims to generate a multi-sentence description of an untrimmed video with several temporal event locations in coherent storytelling. Following the human perception process, where the scene is ef…

DiversitySentenceVideo Captioning