paper-with-me

홈 › Papers

Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval

2025-12-23 · Nguyen Lam Phu Quy, Pham Phu Hoa, Tran Chi Nguyen, Dao Sy Duy Minh, Nguyen Hoang Minh Ngoc, Huynh Trung Kiet arxiv

Real-world image captions often lack contextual depth, omitting crucial details such as event background, temporal cues, outcomes, and named entities that are not visually discernible. This gap limits the effectiveness of image understanding in domains like journalism, education, and digital archives, where richer, more informative descriptions are essential. To address this, we propose a multimodal pipeline that augments visual input with external textual knowledge. Our system retrieves semantically similar images using BEIT-3 (Flickr30k-384 and COCO-384) and SigLIP So-384, reranks them using ORB and SIFT for geometric alignment, and extracts contextual information from related articles via semantic search. A fine-tuned Qwen3 model with QLoRA then integrates this context with base captions generated by Instruct BLIP (Vicuna-7B) to produce event-enriched, context-aware descriptions. Evaluated on the OpenEvents v1 dataset, our approach generates significantly more informative captions compared to traditional methods, showing strong potential for real-world applications requiring deeper visual-textual understanding

📄 PDF Abstract BibTeX arXiv:2512.20042

Code (0)

등록된 구현이 없습니다.

Tasks

Image Captioning

Similar Papers 제목 키워드 기반

Multimodal Arabic Captioning with Interpretable Visual Concept Integration

2025-09-29 · Passant Elchafei, Amany Fashwan arxiv

We present VLCAP, an Arabic image captioning framework that integrates CLIP-based visual label retrieval with multimodal text generation. Rather than relying solely on end-to-end captioning, VLCAP grounds generation in i…

Image CaptioningText Generation

ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding

2025-05-09 · Shuai Wang, Ivona Najdenkoska, Hongyi Zhu, Stevan Rudinac 외

Understanding visual art requires reasoning across multiple perspectives -- cultural, historical, and stylistic -- beyond mere object recognition. While recent multimodal large language models (MLLMs) perform well on gen…

Image CaptioningObject RecognitionRAGRetrieval+1

FuseCap: Leveraging Large Language Models for Enriched Fused Image Captions

2023-05-28 · Noam Rotstein, David Bensaid, Shaked Brody, Roy Ganz 외

The advent of vision-language pre-training techniques enhanced substantial progress in the development of models for image captioning. However, these models frequently produce generic captions and may omit semantically i…

AttributeImage CaptioningLanguage ModellingLarge Language Model+3

ReCap: Event-Aware Image Captioning with Article Retrieval and Semantic Gaussian Normalization

2025-09-01 · Thinh-Phuc Nguyen, Thanh-Hai Nguyen, Gia-Huy Dinh, Lam-Huy Nguyen 외 arxiv

Image captioning systems often produce generic descriptions that fail to capture event-level semantics which are crucial for applications like news reporting and digital archiving. We present ReCap, a novel pipeline for …

Image CaptioningImage Retrieval

Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model

2025-05-29 · Reem AlJunaid, Muzammil Behzad

Generating informative and knowledge-rich image captions remains a challenge for many existing captioning models, which often produce generic descriptions that lack specificity and contextual depth. To address this limit…

Image CaptioningLanguage ModelingLanguage ModellingSpecificity