CapText: Large Language Model-based Caption Generation From Image Context and Description
While deep-learning models have been shown to perform well on image-to-text datasets, it is difficult to use them in practice for captioning images. This is because captions traditionally tend to be context-dependent and offer complementary information about an image, while models tend to produce descriptions that describe the visual features of the image. Prior research in caption generation has explored the use of models that generate captions when provided with the images alongside their respective descriptions or contexts. We propose and evaluate a new approach, which leverages existing large language models to generate captions from textual descriptions and context alone, without ever processing the image directly. We demonstrate that after fine-tuning, our approach outperforms current state-of-the-art image-text alignment models like OSCAR-VinVL on this task on the CIDEr metric.
Code (0)
등록된 구현이 없습니다.
Tasks
Caption GenerationImage to textLanguage ModelingLanguage ModellingLarge Language ModelSimilar Papers 제목 키워드 기반
Towards Automatic Satellite Images Captions Generation Using Large Language Models
Automatic image captioning is a promising technique for conveying visual information using natural language. It can benefit various tasks in satellite remote sensing, such as environmental monitoring, resource management…
Image CaptioningManagementNatural Language UnderstandingCross-modal Language Generation using Pivot Stabilization for Web-scale Language Coverage
Cross-modal language generation tasks such as image captioning are directly hurt in their ability to support non-English languages by the trend of data-hungry models combined with the lack of non-English annotations. We …
Image CaptioningText GenerationTranslationSTAIR Captions: Constructing a Large-Scale Japanese Image Caption Dataset
In recent years, automatic generation of image descriptions (captions), that is, image captioning, has attracted a great deal of attention. In this paper, we particularly consider generating Japanese captions for images.…
Image CaptioningMachine TranslationTranslationCONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning
Multilingual vision-language models have made significant strides in image captioning, yet they still lag behind their English counterparts due to limited multilingual training data and costly large-scale model parameter…
Image CaptioningVector Learning for Cross Domain Representations
Recently, generative adversarial networks have gained a lot of popularity for image generation tasks. However, such models are associated with complex learning mechanisms and demand very large relevant datasets. This wor…
DecoderImage CaptioningImage GenerationSentence+1