Image to Language Understanding: Captioning approach
Extracting context from visual representations is of utmost importance in the advancement of Computer Science. Representation of such a format in Natural Language has a huge variety of applications such as helping the visually impaired etc. Such an approach is a combination of Computer Vision and Natural Language techniques which is a hard problem to solve. This project aims to compare different approaches for solving the image captioning problem. In specific, the focus was on comparing two different types of models: Encoder-Decoder approach and a Multi-model approach. In the encoder-decoder approach, inject and merge architectures were compared against a multi-modal image captioning approach based primarily on object detection. These approaches have been compared on the basis on state of the art sentence comparison metrics such as BLEU, GLEU, Meteor, and Rouge on a subset of the Google Conceptual captions dataset which contains 100k images. On the basis of this comparison, we observed that the best model was the Inception injected encoder model. This best approach has been deployed as a web-based system. On uploading an image, such a system will output the best caption associated with the image.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderImage Captioningobject-detectionObject DetectionSentenceSimilar Papers 제목 키워드 기반
A Thorough Review on Recent Deep Learning Methodologies for Image Captioning
Image Captioning is a task that combines computer vision and natural language processing, where it aims to generate descriptive legends for images. It is a two-fold process relying on accurate image understanding and cor…
Caption GenerationDescriptiveImage CaptioningMeta-LearningFine-Grained Video Captioning through Scene Graph Consolidation
Recent advances in visual language models (VLMs) have significantly improved image captioning, but extending these gains to video understanding remains challenging due to the scarcity of fine-grained video captioning dat…
Caption GenerationImage CaptioningVideo CaptioningVideo UnderstandingImage Captioning based on Deep Reinforcement Learning
Recently it has shown that the policy-gradient methods for reinforcement learning have been utilized to train deep end-to-end systems on natural language processing tasks. What's more, with the complexity of understandin…
Deep Reinforcement LearningImage CaptioningPolicy Gradient Methodsreinforcement-learning+2Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
Image captioning involves generating textual descriptions from input images, bridging the gap between computer vision and natural language processing. Recent advancements in transformer-based models have significantly im…
Caption GenerationImage CaptioningScene UnderstandingSurveyMulti-Level Policy and Reward Reinforcement Learning for Image Captioning
Image captioning is one of the most challenging hallmarks of AI, due to its complexity in visual and natural language understanding. As it is essentially a sequential prediction task, recent advances in image captioning …
Image CaptioningNatural Language Understandingreinforcement-learningReinforcement Learning+2