paper-with-me

Papers

CapText: Large Language Model-based Caption Generation From Image Context and Description

2023-06-01 · Shinjini Ghosh, Sagnik Anupam

While deep-learning models have been shown to perform well on image-to-text datasets, it is difficult to use them in practice for captioning images. This is because captions traditionally tend to be context-dependent and offer complementary information about an image, while models tend to produce descriptions that describe the visual features of the image. Prior research in caption generation has explored the use of models that generate captions when provided with the images alongside their respective descriptions or contexts. We propose and evaluate a new approach, which leverages existing large language models to generate captions from textual descriptions and context alone, without ever processing the image directly. We demonstrate that after fine-tuning, our approach outperforms current state-of-the-art image-text alignment models like OSCAR-VinVL on this task on the CIDEr metric.

📄 PDF Abstract BibTeX arXiv:2306.00301

Code (0)

등록된 구현이 없습니다.

Tasks

Caption GenerationImage to textLanguage ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

Towards Automatic Satellite Images Captions Generation Using Large Language Models

2023-10-17 · Yingxu He, Qiqi Sun

Automatic image captioning is a promising technique for conveying visual information using natural language. It can benefit various tasks in satellite remote sensing, such as environmental monitoring, resource management…

Image CaptioningManagementNatural Language Understanding

Cross-modal Language Generation using Pivot Stabilization for Web-scale Language Coverage

2020-05-01 · ACL 2020 6 · Ashish V. Thapliyal, Radu Soricut

Cross-modal language generation tasks such as image captioning are directly hurt in their ability to support non-English languages by the trend of data-hungry models combined with the lack of non-English annotations. We …

Image CaptioningText GenerationTranslation

STAIR Captions: Constructing a Large-Scale Japanese Image Caption Dataset

2017-05-02 · ACL 2017 7 · Yuya Yoshikawa, Yutaro Shigeto, Akikazu Takeuchi

In recent years, automatic generation of image descriptions (captions), that is, image captioning, has attracted a great deal of attention. In this paper, we particularly consider generating Japanese captions for images.…

Image CaptioningMachine TranslationTranslation

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning

2025-07-27 · George Ibrahim, Rita Ramos, Yova Kementchedjhieva arxiv

Multilingual vision-language models have made significant strides in image captioning, yet they still lag behind their English counterparts due to limited multilingual training data and costly large-scale model parameter…

Image Captioning

Vector Learning for Cross Domain Representations

2018-09-27 · Shagan Sah, Chi Zhang, Thang Nguyen, Dheeraj Kumar Peri 외

Recently, generative adversarial networks have gained a lot of popularity for image generation tasks. However, such models are associated with complex learning mechanisms and demand very large relevant datasets. This wor…

DecoderImage CaptioningImage GenerationSentence+1