paper-with-me

Papers

A Thorough Review on Recent Deep Learning Methodologies for Image Captioning

2021-07-28 · Ahmed Elhagry, Karima Kadaoui

Image Captioning is a task that combines computer vision and natural language processing, where it aims to generate descriptive legends for images. It is a two-fold process relying on accurate image understanding and correct language understanding both syntactically and semantically. It is becoming increasingly difficult to keep up with the latest research and findings in the field of image captioning due to the growing amount of knowledge available on the topic. There is not, however, enough coverage of those findings in the available review papers. We perform in this paper a run-through of the current techniques, datasets, benchmarks and evaluation metrics used in image captioning. The current research on the field is mostly focused on deep learning-based methods, where attention mechanisms along with deep reinforcement and adversarial learning appear to be in the forefront of this research topic. In this paper, we review recent methodologies such as UpDown, OSCAR, VIVO, Meta Learning and a model that uses conditional generative adversarial nets. Although the GAN-based model achieves the highest score, UpDown represents an important basis for image captioning and OSCAR and VIVO are more useful as they use novel object captioning. This review paper serves as a roadmap for researchers to keep up to date with the latest contributions made in the field of image caption generation.

📄 PDF Abstract BibTeX arXiv:2107.13114

Code (0)

등록된 구현이 없습니다.

Tasks

Caption GenerationDescriptiveImage CaptioningMeta-Learning

Methods 이 논문이 사용한 방법론

OSCAR OSCAR is a new learning method that uses object tags detected in images as anchor points to ease the learning of image-text alignment. The model take a triple as input…

Similar Papers 제목 키워드 기반

Pixels to Prose: Understanding the art of Image Captioning

2024-08-28 · Hrishikesh Singh, Aarti Sharma, Millie Pant

In the era of evolving artificial intelligence, machines are increasingly emulating human-like capabilities, including visual perception and linguistic expression. Image captioning stands at the intersection of these dom…

DescriptiveImage CaptioningNavigate

A Review of Methodologies for Natural-Language-Facilitated Human-Robot Cooperation

2017-01-30 · Rui Liu, Xiaoli Zhang

Natural-language-facilitated human-robot cooperation (NLC) refers to using natural language (NL) to facilitate interactive information sharing and task executions with a common goal constraint between robots and humans. …

Autonomous Navigation

A Comprehensive Review for MRF and CRF Approaches in Pathology Image Analysis

2020-09-29 · Yixin Li, Chen Li, Xiaoyan Li, Kai Wang 외

Pathology image analysis is an essential procedure for clinical diagnosis of many diseases. To boost the accuracy and objectivity of detection, nowadays, an increasing number of computer-aided diagnosis (CAD) system is p…

Neural Attention for Image Captioning: Review of Outstanding Methods

2021-11-29 · Zanyar Zohourianshahzadi, Jugal K. Kalita

Image captioning is the task of automatically generating sentences that describe an input image in the best way possible. The most successful techniques for automatically generating image captions have recently used atte…

DecoderDeep LearningImage Captioning

Advancing Image Super-resolution Techniques in Remote Sensing: A Comprehensive Survey

2025-05-29 · Yunliang Qi, Meng Lou, Yimin Liu, Lu Li 외

Remote sensing image super-resolution (RSISR) is a crucial task in remote sensing image processing, aiming to reconstruct high-resolution (HR) images from their low-resolution (LR) counterparts. Despite the growing numbe…

Image Super-ResolutionSuper-Resolution