paper-with-me

홈 › Papers

Compressed Image Captioning using CNN-based Encoder-Decoder Framework

2024-04-28 · Md Alif Rahman Ridoy, M Mahmud Hasan, Shovon Bhowmick

In today's world, image processing plays a crucial role across various fields, from scientific research to industrial applications. But one particularly exciting application is image captioning. The potential impact of effective image captioning is vast. It can significantly boost the accuracy of search engines, making it easier to find relevant information. Moreover, it can greatly enhance accessibility for visually impaired individuals, providing them with a more immersive experience of digital content. However, despite its promise, image captioning presents several challenges. One major hurdle is extracting meaningful visual information from images and transforming it into coherent language. This requires bridging the gap between the visual and linguistic domains, a task that demands sophisticated algorithms and models. Our project is focused on addressing these challenges by developing an automatic image captioning architecture that combines the strengths of convolutional neural networks (CNNs) and encoder-decoder models. The CNN model is used to extract the visual features from images, and later, with the help of the encoder-decoder framework, captions are generated. We also did a performance comparison where we delved into the realm of pre-trained CNN models, experimenting with multiple architectures to understand their performance variations. In our quest for optimization, we also explored the integration of frequency regularization techniques to compress the "AlexNet" and "EfficientNetB0" model. We aimed to see if this compressed model could maintain its effectiveness in generating image captions while being more resource-efficient.

📄 PDF Abstract BibTeX arXiv:2404.18062

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderImage Captioning

Similar Papers 제목 키워드 기반

Review Networks for Caption Generation

2016-05-25 · NeurIPS 2016 12 · Zhilin Yang, Ye Yuan, Yuexin Wu, Ruslan Salakhutdinov 외

We propose a novel extension of the encoder-decoder framework, called a review network. The review network is generic and can enhance any existing encoder- decoder model: in this paper, we consider RNN decoders with both…

Caption GenerationDecoderImage Captioning

Recurrent Fusion Network for Image Captioning

2018-07-26 · ECCV 2018 9 · Wenhao Jiang, Lin Ma, Yu-Gang Jiang, Wei Liu 외

Recently, much advance has been made in image captioning, and an encoder-decoder framework has been adopted by all the state-of-the-art models. Under this framework, an input image is encoded by a convolutional neural ne…

DecoderImage Captioning

Learning to Guide Decoding for Image Captioning

2018-04-03 · Wenhao Jiang, Lin Ma, Xinpeng Chen, Hanwang Zhang 외

Recently, much advance has been made in image captioning, and an encoder-decoder framework has achieved outstanding performance for this task. In this paper, we propose an extension of the encoder-decoder framework by ad…

AttributeDecoderImage Captioning

FE-LWS: Refined Image-Text Representations via Decoder Stacking and Fused Encodings for Remote Sensing Image Captioning

2025-02-13 · Swadhin Das, Raksha Sharma

Remote sensing image captioning aims to generate descriptive text from remote sensing images, typically employing an encoder-decoder framework. In this setup, a convolutional neural network (CNN) extracts feature represe…

Caption GenerationDecoderDescriptiveImage Captioning

A Scaled Encoder Decoder Network for Image Captioning in Hindi

2021-12-01 · ICON 2021 12 · Santosh Kumar Mishra, Sriparna Saha, Pushpak Bhattacharyya

Image captioning is a prominent research area in computer vision and natural language processing, which automatically generates natural language descriptions for images. Most of the existing works have focused on develop…

DecoderDeep LearningImage Captioning