Review Networks for Caption Generation
We propose a novel extension of the encoder-decoder framework, called a review network. The review network is generic and can enhance any existing encoder- decoder model: in this paper, we consider RNN decoders with both CNN and RNN encoders. The review network performs a number of review steps with attention mechanism on the encoder hidden states, and outputs a thought vector after each review step; the thought vectors are used as the input of the attention mechanism in the decoder. We show that conventional encoder-decoders are a special case of our framework. Empirically, we show that our framework improves over state-of- the-art encoder-decoder systems on the tasks of image captioning and source code captioning.
Code (0)
등록된 구현이 없습니다.
Tasks
Caption GenerationDecoderImage CaptioningSimilar Papers 제목 키워드 기반
A Thorough Review on Recent Deep Learning Methodologies for Image Captioning
Image Captioning is a task that combines computer vision and natural language processing, where it aims to generate descriptive legends for images. It is a two-fold process relying on accurate image understanding and cor…
Caption GenerationDescriptiveImage CaptioningMeta-LearningImage Based Review Text Generation with Emotional Guidance
In the current field of computer vision, automatically generating texts from given images has been a fully worked technique. Up till now, most works of this area focus on image content describing, namely image-captioning…
Image CaptioningText GenerationDeep Learning Approaches on Image Captioning: A Review
Image captioning is a research area of immense importance, aiming to generate natural language descriptions for visual content in the form of still images. The advent of deep learning and more recently vision-language pr…
Caption GenerationDeep LearningHallucinationImage Captioning+1Retrieval-Guided Generation for Safer Histopathology Image Captioning
Generative vision-language models can produce fluent medical image captions but remain prone to hallucination, over-specific diagnostic claims, and factual inconsistency-serious issues in pathology. We investigate retrie…
Image CaptioningA Review of Deep Learning for Video Captioning
Video captioning (VC) is a fast-moving, cross-disciplinary area of research that bridges work in the fields of computer vision, natural language processing (NLP), linguistics, and human-computer interaction. In essence, …
Deep LearningDense Video CaptioningQuestion AnsweringRetrieval+3