Word to Sentence Visual Semantic Similarity for Caption Generation: Lessons Learned
This paper focuses on enhancing the captions generated by image-caption generation systems. We propose an approach for improving caption generation systems by choosing the most closely related output to the image rather than the most likely output produced by the model. Our model revises the language generation output beam search from a visual context perspective. We employ a visual semantic measure in a word and sentence level manner to match the proper caption to the related information in the image. The proposed approach can be applied to any caption system as a post-processing based method.
Code (0)
등록된 구현이 없습니다.
Tasks
Caption GenerationSemantic SimilaritySemantic Textual SimilaritySentenceText GenerationSimilar Papers 제목 키워드 기반
Learning semantic sentence representations from visually grounded language without lexical knowledge
Current approaches to learning semantic representations of sentences often use prior word-level knowledge. The current study aims to leverage visual information in order to capture sentence level semantics without the ne…
Grounded language learningLearning Semantic RepresentationsRetrievalSemantic Similarity+5Contrastive Visual Semantic Pretraining Magnifies the Semantics of Natural Language Representations
We examine the effects of contrastive visual semantic pretraining by comparing the geometry and semantic properties of contextualized English language representations formed by GPT-2 and CLIP, a zero-shot multimodal imag…
Image CaptioningSemantic Textual SimilaritySentenceSentence Embeddings+1From Captions to Visual Concepts and Back
This paper presents a novel approach for automatically generating image descriptions: visual detectors, language models, and multimodal similarity models learnt directly from a dataset of image captions. We use multiple …
Image CaptioningLanguage ModelingLanguage ModellingMultiple Instance Learning+2Comprehending and Ordering Semantics for Image Captioning
Comprehending the rich semantics in an image and ordering them in linguistic order are essential to compose a visually-grounded and linguistically coherent description for image captioning. Modern techniques commonly cap…
Cross-Modal RetrievalImage CaptioningRetrievalSentenceShow, Tell and Summarize: Dense Video Captioning Using Visual Cue Aided Sentence Summarization
In this work, we propose a division-and-summarization (DaS) framework for dense video captioning. After partitioning each untrimmed long video as multiple event proposals, where each event proposal consists of a set of s…
Dense Video CaptioningDescriptiveSentenceSentence Summarization+1