DIFNet: Boosting Visual Information Flow for Image Captioning
Current Image captioning (IC) methods predict textual words sequentially based on the input visual information from the visual feature extractor and the partially generated sentence information. However, for most cases, the partially generated sentence may dominate the target word prediction due to the insufficiency of visual information, making the generated descriptions irrelevant to the content of the given image. In this paper, we propose a Dual Information Flow Network (DIFNet) to address this issue, which takes segmentation feature as another visual information source to enhance the contribution of visual information for prediction. To maximize the use of two information flows, we also propose an effective feature fusion module termed Iterative Independent Layer Normalization (IILN) which can condense the most relevant inputs while retraining modality-specific information in each flow. Experiments show that our method is able to enhance the dependence of prediction on visual information, making word prediction more focused on the visual content, and thus achieve new state-of-the-art performance on the MSCOCO dataset, e.g., 136.2 CIDEr on COCO Karpathy test split.
Code (0)
등록된 구현이 없습니다.
Tasks
Image CaptioningPredictionSentenceMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Get Rid of Suspended Animation Problem: Deep Diffusive Neural Network on Graph Semi-Supervised Classification
Existing graph neural networks may suffer from the "suspended animation problem" when the model architecture goes deep. Meanwhile, for some graph learning scenarios, e.g., nodes with text/image attributes or graphs with …
General ClassificationGraph LearningGraph Neural NetworkGraph Representation Learning+2DifNet: Semantic Segmentation by Diffusion Networks
Deep Neural Networks (DNNs) have recently shown state of the art performance on semantic segmentation tasks, however, they still suffer from problems of poor boundary localization and spatial fragmented predictions. The …
SegmentationSemantic SegmentationBoosting Dermatoscopic Lesion Segmentation via Diffusion Models with Visual and Textual Prompts
Image synthesis approaches, e.g., generative adversarial networks, have been popular as a form of data augmentation in medical image analysis tasks. It is primarily beneficial to overcome the shortage of publicly accessi…
Data AugmentationImage GenerationLesion SegmentationMedical Image Analysis+1PointMCD: Boosting Deep Point Cloud Encoders via Multi-view Cross-modal Distillation for 3D Shape Recognition
As two fundamental representation modalities of 3D objects, 3D point clouds and multi-view 2D images record shape information from different domains of geometric structures and visual appearances. In the current deep lea…
3D Shape Classification3D Shape RecognitionTransfer LearningA Deep Features-Based Approach Using Modified ResNet50 and Gradient Boosting for Visual Sentiments Classification
The versatile nature of Visual Sentiment Analysis (VSA) is one reason for its rising profile. It isn't easy to efficiently manage social media data with visual information since previous research has concentrated on Sent…
Sentiment Analysis