paper-with-me

홈 › Papers

Image Captioning based on Deep Reinforcement Learning

2018-09-13 · Haichao Shi, Peng Li, Bo wang, Zhenyu Wang

Recently it has shown that the policy-gradient methods for reinforcement learning have been utilized to train deep end-to-end systems on natural language processing tasks. What's more, with the complexity of understanding image content and diverse ways of describing image content in natural language, image captioning has been a challenging problem to deal with. To the best of our knowledge, most state-of-the-art methods follow a pattern of sequential model, such as recurrent neural networks (RNN). However, in this paper, we propose a novel architecture for image captioning with deep reinforcement learning to optimize image captioning tasks. We utilize two networks called "policy network" and "value network" to collaboratively generate the captions of images. The experiments are conducted on Microsoft COCO dataset, and the experimental results have verified the effectiveness of the proposed method.

📄 PDF Abstract BibTeX arXiv:1809.04835

Code (0)

등록된 구현이 없습니다.

Tasks

Deep Reinforcement LearningImage CaptioningPolicy Gradient Methodsreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Actor-Critic Sequence Training for Image Captioning

2017-06-29 · Li Zhang, Flood Sung, Feng Liu, Tao Xiang 외

Generating natural language descriptions of images is an important capability for a robot or other visual-intelligence driven AI agent that may need to communicate with human users about what it is seeing. Such image cap…

AI AgentImage Captioningreinforcement-learningReinforcement Learning+1

Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning

2026-08-21 · Haonan Jia, Shichao Dong, Zenghui Sun, Jiawen Zheng 외 arxiv

Reinforcement Learning (RL) has demonstrated significant gains in image captioning, yet it is still limited in encouraging Large Vision-Language Models (LVLMs) to explore novel reasoning strategies. This limitation leads…

Reinforcement LearningImage Captioning

Multi-Level Policy and Reward Reinforcement Learning for Image Captioning

2018-06-15 · IJCAI 2018 6 · An-An Liu1, Ning Xu1, Hanwang Zhang2, Weizhi Nie1 외

Image captioning is one of the most challenging hallmarks of AI, due to its complexity in visual and natural language understanding. As it is essentially a sequential prediction task, recent advances in image captioning …

Image CaptioningNatural Language Understandingreinforcement-learningReinforcement Learning+2

Contextualized Keyword Representations for Multi-modal Retinal Image Captioning

2021-04-26 · Jia-Hong Huang, Ting-Wei Wu, Marcel Worring

Medical image captioning automatically generates a medical description to describe the content of a given medical image. A traditional medical image captioning model creates a medical description only based on a single m…

AvgImage Captioning

Multi-modal reward for visual relationships-based image captioning

2023-03-19 · Ali Abedi, Hossein Karshenas, Peyman Adibi

Deep neural networks have achieved promising results in automatic image captioning due to their effective representation learning and context-based content generation capabilities. As a prominent type of deep features us…

Caption GenerationDeep Reinforcement LearningImage CaptioningModel Optimization+3