Generating Image Descriptions via Sequential Cross-Modal Alignment Guided by Human Gaze
When speakers describe an image, they tend to look at objects before mentioning them. In this paper, we investigate such sequential cross-modal alignment by modelling the image description generation process computationally. We take as our starting point a state-of-the-art image captioning system and develop several model variants that exploit information from human gaze patterns recorded during language production. In particular, we propose the first approach to image description generation where visual processing is modelled $\textit{sequentially}$. Our experiments and analyses confirm that better descriptions can be obtained by exploiting gaze-driven attention and shed light on human cognitive processes by comparing different ways of aligning the gaze modality with language production. We find that processing gaze data sequentially leads to descriptions that are better aligned to those produced by speakers, more diverse, and more natural${-}$particularly when gaze is encoded with a dedicated recurrent component.
Code (1)
Tasks
cross-modal alignmentImage CaptioningImage DescriptionSimilar Papers 제목 키워드 기반
Generating Descriptions for Sequential Images with Local-Object Attention and Global Semantic Context Modelling
In this paper, we propose an end-to-end CNN-LSTM model for generating descriptions for sequential images with a local-object attention mechanism. To generate coherent descriptions, we capture global semantic context usin…
Trajectory-aware Cross-view Geo-localization with Sequential Observations
Cross-view geo-localization matches ground-level observations against geo-tagged satellite imagery. Recent methods show that sequential queries such as video clips yield richer spatiotemporal cues than single images, yet…
CapEnrich: Enriching Caption Semantics for Web Images via Cross-modal Pre-trained Knowledge
Automatically generating textual descriptions for massive unlabeled images on the web can greatly benefit realistic web applications, e.g. multimodal retrieval and recommendation. However, existing models suffer from the…
Concept AlignmentRetrievalDeep Visual-Semantic Alignments for Generating Image Descriptions
We present a model that generates natural language descriptions of images and their regions. Our approach leverages datasets of images and their sentence descriptions to learn about the inter-modal correspondences betwee…
Cross-Modal RetrievalImage CaptioningImage-to-Text RetrievalRetrieval+1Multimodal RAG Enhanced Visual Description
Textual descriptions for multimodal inputs entail recurrent refinement of queries to produce relevant output images. Despite efforts to address challenges such as scaling model size and data volume, the cost associated w…