paper-with-me

Papers

Generating Image Descriptions via Sequential Cross-Modal Alignment Guided by Human Gaze

2020-11-09 · EMNLP 2020 11 · Ece Takmaz, Sandro Pezzelle, Lisa Beinborn, Raquel Fernández

When speakers describe an image, they tend to look at objects before mentioning them. In this paper, we investigate such sequential cross-modal alignment by modelling the image description generation process computationally. We take as our starting point a state-of-the-art image captioning system and develop several model variants that exploit information from human gaze patterns recorded during language production. In particular, we propose the first approach to image description generation where visual processing is modelled $\textit{sequentially}$. Our experiments and analyses confirm that better descriptions can be obtained by exploiting gaze-driven attention and shed light on human cognitive processes by comparing different ways of aligning the gaze modality with language production. We find that processing gaze data sequentially leads to descriptions that are better aligned to those produced by speakers, more diverse, and more natural${-}$particularly when gaze is encoded with a dedicated recurrent component.

📄 PDF Abstract BibTeX arXiv:2011.04592

Code (1)

dmg-illc/didec-seq-gen 공식 구현 pytorch

Tasks

cross-modal alignmentImage CaptioningImage Description

Similar Papers 제목 키워드 기반

Generating Descriptions for Sequential Images with Local-Object Attention and Global Semantic Context Modelling

2020-12-02 · Jing Su, Chenghua Lin, Mian Zhou, Qingyun Dai 외

In this paper, we propose an end-to-end CNN-LSTM model for generating descriptions for sequential images with a local-object attention mechanism. To generate coherent descriptions, we capture global semantic context usin…

Trajectory-aware Cross-view Geo-localization with Sequential Observations

2026-07-16 · Tianyi Gao, Jiayu Lin, Danielle Beaulieu, Nathan Jacobs arxiv

Cross-view geo-localization matches ground-level observations against geo-tagged satellite imagery. Recent methods show that sequential queries such as video clips yield richer spatiotemporal cues than single images, yet…

CapEnrich: Enriching Caption Semantics for Web Images via Cross-modal Pre-trained Knowledge

2022-11-17 · Linli Yao, Weijing Chen, Qin Jin

Automatically generating textual descriptions for massive unlabeled images on the web can greatly benefit realistic web applications, e.g. multimodal retrieval and recommendation. However, existing models suffer from the…

Concept AlignmentRetrieval

Deep Visual-Semantic Alignments for Generating Image Descriptions

2014-12-07 · CVPR 2015 6 · Andrej Karpathy, Li Fei-Fei

We present a model that generates natural language descriptions of images and their regions. Our approach leverages datasets of images and their sentence descriptions to learn about the inter-modal correspondences betwee…

Cross-Modal RetrievalImage CaptioningImage-to-Text RetrievalRetrieval+1

Multimodal RAG Enhanced Visual Description

2025-08-06 · Amit Kumar Jaiswal, Haiming Liu, Ingo Frommholz arxiv

Textual descriptions for multimodal inputs entail recurrent refinement of queries to produce relevant output images. Despite efforts to address challenges such as scaling model size and data volume, the cost associated w…