paper-with-me

홈 › Papers

New Encoder Learning for Captioning Heavy Rain Images via Semantic Visual Feature Matching

2021-05-28 · Chang-Hwan Son, Pung-Hwi Ye

Image captioning generates text that describes scenes from input images. It has been developed for high quality images taken in clear weather. However, in bad weather conditions, such as heavy rain, snow, and dense fog, the poor visibility owing to rain streaks, rain accumulation, and snowflakes causes a serious degradation of image quality. This hinders the extraction of useful visual features and results in deteriorated image captioning performance. To address practical issues, this study introduces a new encoder for captioning heavy rain images. The central idea is to transform output features extracted from heavy rain input images into semantic visual features associated with words and sentence context. To achieve this, a target encoder is initially trained in an encoder-decoder framework to associate visual features with semantic words. Subsequently, the objects in a heavy rain image are rendered visible by using an initial reconstruction subnetwork (IRS) based on a heavy rain model. The IRS is then combined with another semantic visual feature matching subnetwork (SVFMS) to match the output features of the IRS with the semantic visual features of the pretrained target encoder. The proposed encoder is based on the joint learning of the IRS and SVFMS. It is is trained in an end-to-end manner, and then connected to the pretrained decoder for image captioning. It is experimentally demonstrated that the proposed encoder can generate semantic visual features associated with words even from heavy rain images, thereby increasing the accuracy of the generated captions.

📄 PDF Abstract BibTeX arXiv:2105.13753

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderImage CaptioningSentence

Similar Papers 제목 키워드 기반

Relational Reasoning using Prior Knowledge for Visual Captioning

2019-06-04 · Jingyi Hou, Xinxiao Wu, Yayun Qi, Wentian Zhao 외

Exploiting relationships among objects has achieved remarkable progress in interpreting images or videos by natural language. Most existing methods resort to first detecting objects and their relationships, and then gene…

Image Captioningobject-detectionObject DetectionRelational Reasoning+2

An Image captioning algorithm based on the Hybrid Deep Learning Technique (CNN+GRU)

2023-01-06 · Rana Adnan Ahmad, Muhammad Azhar, Hina Sattar

Image captioning by the encoder-decoder framework has shown tremendous advancement in the last decade where CNN is mainly used as encoder and LSTM is used as a decoder. Despite such an impressive achievement in terms of …

DecoderImage Captioning

Finding It at Another Side: A Viewpoint-Adapted Matching Encoder for Change Captioning

2020-09-30 · ECCV 2020 8 · Xiangxi Shi, Xu Yang, Jiuxiang Gu, Shafiq Joty 외

Change Captioning is a task that aims to describe the difference between images with natural language. Most existing methods treat this problem as a difference judgment without the existence of distractors, such as viewp…

Reinforcement Learning (RL)

Change Captioning in Remote Sensing: Evolution to SAT-Cap -- A Single-Stage Transformer Approach

2025-01-14 · Yuduo Wang, Weikang Yu, Pedram Ghamisi

Change captioning has become essential for accurately describing changes in multi-temporal remote sensing data, providing an intuitive way to monitor Earth's dynamics through natural language. However, existing change ca…

A TextGCN-Based Decoding Approach for Improving Remote Sensing Image Captioning

2024-09-27 · Swadhin Das, Raksha Sharma

Remote sensing images are highly valued for their ability to address complex real-world issues such as risk management, security, and meteorology. However, manually captioning these images is challenging and requires spe…

DecoderFairnessImage CaptioningSentence