paper-with-me

홈 › Papers

Word to Sentence Visual Semantic Similarity for Caption Generation: Lessons Learned

2022-09-26 · Ahmed Sabir

This paper focuses on enhancing the captions generated by image-caption generation systems. We propose an approach for improving caption generation systems by choosing the most closely related output to the image rather than the most likely output produced by the model. Our model revises the language generation output beam search from a visual context perspective. We employ a visual semantic measure in a word and sentence level manner to match the proper caption to the related information in the image. The proposed approach can be applied to any caption system as a post-processing based method.

📄 PDF Abstract BibTeX arXiv:2209.12817

Code (0)

등록된 구현이 없습니다.

Tasks

Caption GenerationSemantic SimilaritySemantic Textual SimilaritySentenceText Generation

Similar Papers 제목 키워드 기반

Learning semantic sentence representations from visually grounded language without lexical knowledge

2019-03-27 · Danny Merkx, Stefan Frank

Current approaches to learning semantic representations of sentences often use prior word-level knowledge. The current study aims to leverage visual information in order to capture sentence level semantics without the ne…

Grounded language learningLearning Semantic RepresentationsRetrievalSemantic Similarity+5

Contrastive Visual Semantic Pretraining Magnifies the Semantics of Natural Language Representations

2022-03-14 · ACL 2022 5 · Robert Wolfe, Aylin Caliskan

We examine the effects of contrastive visual semantic pretraining by comparing the geometry and semantic properties of contextualized English language representations formed by GPT-2 and CLIP, a zero-shot multimodal imag…

Image CaptioningSemantic Textual SimilaritySentenceSentence Embeddings+1

From Captions to Visual Concepts and Back

2014-11-18 · CVPR 2015 6 · Hao Fang, Saurabh Gupta, Forrest Iandola, Rupesh Srivastava 외

This paper presents a novel approach for automatically generating image descriptions: visual detectors, language models, and multimodal similarity models learnt directly from a dataset of image captions. We use multiple …

Image CaptioningLanguage ModelingLanguage ModellingMultiple Instance Learning+2

Comprehending and Ordering Semantics for Image Captioning

2022-06-14 · CVPR 2022 1 · Yehao Li, Yingwei Pan, Ting Yao, Tao Mei

Comprehending the rich semantics in an image and ordering them in linguistic order are essential to compose a visually-grounded and linguistically coherent description for image captioning. Modern techniques commonly cap…

Cross-Modal RetrievalImage CaptioningRetrievalSentence

Show, Tell and Summarize: Dense Video Captioning Using Visual Cue Aided Sentence Summarization

2025-06-25 · Zhiwang Zhang, Dong Xu, Wanli Ouyang, Chuanqi Tan

In this work, we propose a division-and-summarization (DaS) framework for dense video captioning. After partitioning each untrimmed long video as multiple event proposals, where each event proposal consists of a set of s…

Dense Video CaptioningDescriptiveSentenceSentence Summarization+1