paper-with-me

홈 › Papers

SmallCap: Lightweight Image Captioning Prompted with Retrieval Augmentation

2022-09-30 · CVPR 2023 1 · Rita Ramos, Bruno Martins, Desmond Elliott, Yova Kementchedjhieva

Recent advances in image captioning have focused on scaling the data and model size, substantially increasing the cost of pre-training and finetuning. As an alternative to large models, we present SmallCap, which generates a caption conditioned on an input image and related captions retrieved from a datastore. Our model is lightweight and fast to train, as the only learned parameters are in newly introduced cross-attention layers between a pre-trained CLIP encoder and GPT-2 decoder. SmallCap can transfer to new domains without additional finetuning and can exploit large-scale data in a training-free fashion since the contents of the datastore can be readily replaced. Our experiments show that SmallCap, trained only on COCO, has competitive performance on this benchmark, and also transfers to other domains without retraining, solely through retrieval from target-domain data. Further improvement is achieved through the training-free exploitation of diverse human-labeled and web data, which proves to be effective for a range of domains, including the nocaps benchmark, designed to test generalization to unseen visual concepts.

📄 PDF Abstract BibTeX arXiv:2209.15323

Code (1)

ritaramo/smallcap 공식 구현 pytorch

Tasks

DecoderImage CaptioningRetrieval

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Test 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Multi-Head Attention 설명 없음
Adam 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning

2024-06-04 · Wenyan Li, Jiaang Li, Rita Ramos, Raphael Tang 외

Recent advances in retrieval-augmented models for image captioning highlight the benefit of retrieving related captions for efficient, lightweight models with strong domain-transfer capabilities. While these models demon…

Image CaptioningRetrieval

RACap: Relation-Aware Prompting for Lightweight Retrieval-Augmented Image Captioning

2025-09-19 · Xiaosheng Long, Hanyu Wang, Zhentao Song, Kun Luo 외 arxiv

Recent retrieval-augmented image captioning methods incorporate external knowledge to compensate for the limitations in comprehending complex scenes. However, current approaches face challenges in relation modeling: (1) …

Image Captioning

DualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual Prompts

2025-10-28 · Binbin Li, Guimiao Yang, Zisen Qi, Haiping Wang 외 arxiv

Recent lightweight retrieval-augmented image caption models often utilize retrieved data solely as text prompts, thereby creating a semantic gap by leaving the original visual features unenhanced, particularly for object…

Image-to-Text RetrievalImage CaptioningImage Retrieval

ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning

2024-12-26 · Taewhan Kim, Soeun Lee, Si-Woo Kim, Dong-Jin Kim

Recent lightweight image captioning models using retrieved data mainly focus on text prompts. However, previous works only utilize the retrieved text as text prompts, and the visual information relies only on the CLIP vi…

Image CaptioningRetrieval

PromptCap: Prompt-Guided Task-Aware Image Captioning

2022-11-15 · Yushi Hu, Hang Hua, Zhengyuan Yang, Weijia Shi 외

Knowledge-based visual question answering (VQA) involves questions that require world knowledge beyond the image to yield the correct answer. Large language models (LMs) like GPT-3 are particularly helpful for this task …

Image CaptioningLanguage ModellingQuestion AnsweringRetrieval+4