paper-with-me

Papers

RefineCap: Concept-Aware Refinement for Image Captioning

2021-09-08 · Yekun Chai, Shuo Jin, Junliang Xing

Automatically translating images to texts involves image scene understanding and language modeling. In this paper, we propose a novel model, termed RefineCap, that refines the output vocabulary of the language decoder using decoder-guided visual semantics, and implicitly learns the mapping between visual tag words and images. The proposed Visual-Concept Refinement method can allow the generator to attend to semantic details in the image, thereby generating more semantically descriptive captions. Our model achieves superior performance on the MS-COCO dataset in comparison with previous visual-concept based models.

📄 PDF Abstract BibTeX arXiv:2109.03529

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderDescriptiveImage CaptioningLanguage ModelingLanguage ModellingScene UnderstandingTAG

Similar Papers 제목 키워드 기반

CLIP Meets Video Captioning: Concept-Aware Representation Learning Does Matter

2021-11-30 · Bang Yang, Tong Zhang, Yuexian Zou

For video captioning, "pre-training and fine-tuning" has become a de facto paradigm, where ImageNet Pre-training (INP) is usually used to encode the video content, then a task-oriented network is fine-tuned from scratch …

Caption GenerationRepresentation LearningVideo Captioning

Towards Accurate Emotion-Attributed Video Captioning via Fine-grained Emotion-Cause Pair Extraction

2026-06-07 · Weidong Chen, Cheng Ye, Zhendong Mao, Liping Wang 외 arxiv

Emotional Video Captioning (EVC) is a challenging task that aims to generate factually accurate and emotionally rich descriptions for videos. Existing EVC methods leverage holistic visual features to mine global emotiona…

Emotion-Cause Pair ExtractionVideo Captioning

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning

2025-07-27 · George Ibrahim, Rita Ramos, Yova Kementchedjhieva arxiv

Multilingual vision-language models have made significant strides in image captioning, yet they still lag behind their English counterparts due to limited multilingual training data and costly large-scale model parameter…

Image Captioning

CIAN: Multi-Stage Framework for Event-Enriched Image Captioning via Retrieval-Augmented Generation

2026-06-16 · Trinh Thi Thu Hien, Trung-Nghia Le arxiv

Event-enriched image captioning describes not only visible content but also the broader context of events, including timing, location, and participants, capabilities missing in most pixel-bound models. We propose the Con…

Image Captioning

Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning

2026-08-21 · Haonan Jia, Shichao Dong, Zenghui Sun, Jiawen Zheng 외 arxiv

Reinforcement Learning (RL) has demonstrated significant gains in image captioning, yet it is still limited in encouraging Large Vision-Language Models (LVLMs) to explore novel reasoning strategies. This limitation leads…

Reinforcement LearningImage Captioning