paper-with-me

홈 › Papers

Image to Language Understanding: Captioning approach

2020-02-21 · Madhavan Seshadri, Malavika Srikanth, Mikhail Belov

Extracting context from visual representations is of utmost importance in the advancement of Computer Science. Representation of such a format in Natural Language has a huge variety of applications such as helping the visually impaired etc. Such an approach is a combination of Computer Vision and Natural Language techniques which is a hard problem to solve. This project aims to compare different approaches for solving the image captioning problem. In specific, the focus was on comparing two different types of models: Encoder-Decoder approach and a Multi-model approach. In the encoder-decoder approach, inject and merge architectures were compared against a multi-modal image captioning approach based primarily on object detection. These approaches have been compared on the basis on state of the art sentence comparison metrics such as BLEU, GLEU, Meteor, and Rouge on a subset of the Google Conceptual captions dataset which contains 100k images. On the basis of this comparison, we observed that the best model was the Inception injected encoder model. This best approach has been deployed as a web-based system. On uploading an image, such a system will output the best caption associated with the image.

📄 PDF Abstract BibTeX arXiv:2002.09536

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderImage Captioningobject-detectionObject DetectionSentence

Similar Papers 제목 키워드 기반

A Thorough Review on Recent Deep Learning Methodologies for Image Captioning

2021-07-28 · Ahmed Elhagry, Karima Kadaoui

Image Captioning is a task that combines computer vision and natural language processing, where it aims to generate descriptive legends for images. It is a two-fold process relying on accurate image understanding and cor…

Caption GenerationDescriptiveImage CaptioningMeta-Learning

Fine-Grained Video Captioning through Scene Graph Consolidation

2025-02-23 · Sanghyeok Chu, Seonguk Seo, Bohyung Han

Recent advances in visual language models (VLMs) have significantly improved image captioning, but extending these gains to video understanding remains challenging due to the scarcity of fine-grained video captioning dat…

Caption GenerationImage CaptioningVideo CaptioningVideo Understanding

Image Captioning based on Deep Reinforcement Learning

2018-09-13 · Haichao Shi, Peng Li, Bo wang, Zhenyu Wang

Recently it has shown that the policy-gradient methods for reinforcement learning have been utilized to train deep end-to-end systems on natural language processing tasks. What's more, with the complexity of understandin…

Deep Reinforcement LearningImage CaptioningPolicy Gradient Methodsreinforcement-learning+2

Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation

2025-06-03 · Israa A. Albadarneh, Bassam H. Hammo, Omar S. Al-Kadi

Image captioning involves generating textual descriptions from input images, bridging the gap between computer vision and natural language processing. Recent advancements in transformer-based models have significantly im…

Caption GenerationImage CaptioningScene UnderstandingSurvey

Multi-Level Policy and Reward Reinforcement Learning for Image Captioning

2018-06-15 · IJCAI 2018 6 · An-An Liu1, Ning Xu1, Hanwang Zhang2, Weizhi Nie1 외

Image captioning is one of the most challenging hallmarks of AI, due to its complexity in visual and natural language understanding. As it is essentially a sequential prediction task, recent advances in image captioning …

Image CaptioningNatural Language Understandingreinforcement-learningReinforcement Learning+2