RefineCap: Concept-Aware Refinement for Image Captioning
Automatically translating images to texts involves image scene understanding and language modeling. In this paper, we propose a novel model, termed RefineCap, that refines the output vocabulary of the language decoder using decoder-guided visual semantics, and implicitly learns the mapping between visual tag words and images. The proposed Visual-Concept Refinement method can allow the generator to attend to semantic details in the image, thereby generating more semantically descriptive captions. Our model achieves superior performance on the MS-COCO dataset in comparison with previous visual-concept based models.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderDescriptiveImage CaptioningLanguage ModelingLanguage ModellingScene UnderstandingTAGSimilar Papers 제목 키워드 기반
CLIP Meets Video Captioning: Concept-Aware Representation Learning Does Matter
For video captioning, "pre-training and fine-tuning" has become a de facto paradigm, where ImageNet Pre-training (INP) is usually used to encode the video content, then a task-oriented network is fine-tuned from scratch …
Caption GenerationRepresentation LearningVideo CaptioningTowards Accurate Emotion-Attributed Video Captioning via Fine-grained Emotion-Cause Pair Extraction
Emotional Video Captioning (EVC) is a challenging task that aims to generate factually accurate and emotionally rich descriptions for videos. Existing EVC methods leverage holistic visual features to mine global emotiona…
Emotion-Cause Pair ExtractionVideo CaptioningCONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning
Multilingual vision-language models have made significant strides in image captioning, yet they still lag behind their English counterparts due to limited multilingual training data and costly large-scale model parameter…
Image CaptioningCIAN: Multi-Stage Framework for Event-Enriched Image Captioning via Retrieval-Augmented Generation
Event-enriched image captioning describes not only visible content but also the broader context of events, including timing, location, and participants, capabilities missing in most pixel-bound models. We propose the Con…
Image CaptioningRe$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning
Reinforcement Learning (RL) has demonstrated significant gains in image captioning, yet it is still limited in encouraging Large Vision-Language Models (LVLMs) to explore novel reasoning strategies. This limitation leads…
Reinforcement LearningImage Captioning