paper-with-me

홈 › Papers

ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning

2024-12-26 · Taewhan Kim, Soeun Lee, Si-Woo Kim, Dong-Jin Kim

Recent lightweight image captioning models using retrieved data mainly focus on text prompts. However, previous works only utilize the retrieved text as text prompts, and the visual information relies only on the CLIP visual embedding. Because of this issue, there is a limitation that the image descriptions inherent in the prompt are not sufficiently reflected in the visual embedding space. To tackle this issue, we propose ViPCap, a novel retrieval text-based visual prompt for lightweight image captioning. ViPCap leverages the retrieved text with image information as visual prompts to enhance the ability of the model to capture relevant visual information. By mapping text prompts into the CLIP space and generating multiple randomized Gaussian distributions, our method leverages sampling to explore randomly augmented distributions and effectively retrieves the semantic features that contain image information. These retrieved features are integrated into the image and designated as the visual prompt, leading to performance improvements on the datasets such as COCO, Flickr30k, and NoCaps. Experimental results demonstrate that ViPCap significantly outperforms prior lightweight captioning models in efficiency and effectiveness, demonstrating the potential for a plug-and-play solution. The source code is available at https://github.com/taewhankim/VIPCAP.

📄 PDF Abstract BibTeX arXiv:2412.19289

Code (1)

taewhankim/vipcap 공식 구현

Tasks

Image CaptioningRetrieval

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
Focus 설명 없음

Similar Papers 제목 키워드 기반

DualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual Prompts

2025-10-28 · Binbin Li, Guimiao Yang, Zisen Qi, Haiping Wang 외 arxiv

Recent lightweight retrieval-augmented image caption models often utilize retrieved data solely as text prompts, thereby creating a semantic gap by leaving the original visual features unenhanced, particularly for object…

Image-to-Text RetrievalImage CaptioningImage Retrieval

AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs

2025-06-19 · Yuan Zhang, Chun-Kai Fan, Sicheng Yu, Junwen Pan 외 arxiv

Inspired by text prompts in large language models, visual prompts have been explored to enhance the perceptual capabilities of large vision-language models (LVLMs). However, performance tends to saturate under single vis…

RACap: Relation-Aware Prompting for Lightweight Retrieval-Augmented Image Captioning

2025-09-19 · Xiaosheng Long, Hanyu Wang, Zhentao Song, Kun Luo 외 arxiv

Recent retrieval-augmented image captioning methods incorporate external knowledge to compensate for the limitations in comprehending complex scenes. However, current approaches face challenges in relation modeling: (1) …

Image Captioning

Dynamic Tool Dependency Retrieval for Lightweight Function Calling

2025-12-18 · Bhrij Patel, Davide Belli, Amir Jalalirad, Maximilian Arnold 외 arxiv

Function calling agents powered by Large Language Models (LLMs) select external tools to automate complex tasks. On-device agents typically use a retrieval module to select relevant tools, improving performance and reduc…

Computational Efficiency

PQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance Prediction

2024-06-07 · CVPR 2025 1 · Eduard Poesina, Adriana Valentina Costache, Adrian-Gabriel Chifu, Josiane Mothe 외

Text-to-image generation has recently emerged as a viable alternative to text-to-image retrieval, due to the visually impressive results of generative diffusion models. Although query performance prediction is an active …

Image GenerationImage RetrievalInformation RetrievalRetrieval+2