How Do Image Description Systems Describe People? A Targeted Assessment of System Competence in the PEOPLE-domain
Evaluations of image description systems are typically domain-general: generated descriptions for the held-out test images are either compared to a set of reference descriptions (using automated metrics), or rated by human judges on one or more Likert scales (for fluency, overall quality, and other quality criteria). While useful, these evaluations do not tell us anything about the kinds of image descriptions that systems are able to produce. Or, phrased differently, these evaluations do not tell us anything about the cognitive capabilities of image description systems. This paper proposes a different kind of assessment, that is able to quantify the extent to which these systems are able to describe humans. This assessment is based on a manual characterisation (a context-free grammar) of English entity labels in the PEOPLE domain, to determine the range of possible outputs. We examined 9 systems to see what kinds of labels they actually use. We found that these systems only use a small subset of at most 13 different kinds of modifiers (e.g. tall and short modify HEIGHT, sad and happy modify MOOD), but 27 kinds of modifiers are never used. Future research could study these semantic dimensions in more detail.
Code (1)
Tasks
Image DescriptionSimilar Papers 제목 키워드 기반
Exploring the Behavior of Classic REG Algorithms in the Description of Characters in 3D Images
Describing people and characters can be very useful in different contexts, such as computational narrative or image description for the visually impaired. However, a review of the existing literature shows that the autom…
Image DescriptionReferring ExpressionReferring expression generationSurvey+1Personalized Image Descriptions from Attention Sequences
People can view the same image differently: they focus on different regions, objects, and details in varying orders and describe them in distinct linguistic styles. This leads to substantial variability in image descript…
Studying Relationships between Human Gaze, Description, and Computer Vision
We posit that user behavior during natural viewing of images contains an abundance of information about the content of images as well as information related to user intent and user defined content importance. In this pap…
Adapting Descriptions of People to the Point of View of a Moving Observer
This paper addresses the task of generating descriptions of people for an observer that is moving within a scene. As the observer moves, the descriptions of the people around him also change. A referring expression gener…
PositionReferring ExpressionReferring expression generationText GenerationTalking about other people: an endless range of possibilities
Image description datasets, such as Flickr30K and MS COCO, show a high degree of variation in the ways that crowd-workers talk about the world. Although this gives us a rich and diverse collection of data to work with, i…
Image DescriptionText Generation