paper-with-me

홈 › Papers

Dual Attention on Pyramid Feature Maps for Image Captioning

2020-11-02 · Litao Yu, Jian Zhang, Qiang Wu

Generating natural sentences from images is a fundamental learning task for visual-semantic understanding in multimedia. In this paper, we propose to apply dual attention on pyramid image feature maps to fully explore the visual-semantic correlations and improve the quality of generated sentences. Specifically, with the full consideration of the contextual information provided by the hidden state of the RNN controller, the pyramid attention can better localize the visually indicative and semantically consistent regions in images. On the other hand, the contextual information can help re-calibrate the importance of feature components by learning the channel-wise dependencies, to improve the discriminative power of visual features for better content description. We conducted comprehensive experiments on three well-known datasets: Flickr8K, Flickr30K and MS COCO, which achieved impressive results in generating descriptive and smooth natural sentences from images. Using either convolution visual features or more informative bottom-up attention features, our composite captioning model achieves very promising performance in a single-model mode. The proposed pyramid attention and dual attention methods are highly modular, which can be inserted into various image captioning modules to further improve the performance.

📄 PDF Abstract BibTeX arXiv:2011.01385

Code (0)

등록된 구현이 없습니다.

Tasks

DescriptiveImage Captioning

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

Pyramid Feature Attention Network for Monocular Depth Prediction

2024-03-03 · Yifang Xu, Chenglei Peng, Ming Li, Yang Li 외

Deep convolutional neural networks (DCNNs) have achieved great success in monocular depth estimation (MDE). However, few existing works take the contributions for MDE of different levels feature maps into account, leadin…

Depth EstimationDepth PredictionMonocular Depth EstimationPrediction

Learning Spatial Pyramid Attentive Pooling in Image Synthesis and Image-to-Image Translation

2019-01-18 · Wei Sun, Tianfu Wu

Image synthesis and image-to-image translation are two important generative learning tasks. Remarkable progress has been made by learning Generative Adversarial Networks (GANs)~\cite{goodfellow2014generative} and cycle-c…

Image GenerationImage-to-Image TranslationTranslation

Interaction-aware Spatio-temporal Pyramid Attention Networks for Action Classification

2018-08-03 · ECCV 2018 9 · Yang Du, Chunfeng Yuan, Bing Li, Lili Zhao 외

Local features at neighboring spatial positions in feature maps have high correlation since their receptive fields are often overlapped. Self-attention usually uses the weighted sum (or other functions) with internal ele…

Action ClassificationClassificationGeneral Classification

Chinese Herbal Recognition based on Competitive Attentional Fusion of Multi-hierarchies Pyramid Features

2018-12-23 · Yingxue Xu, Guihua Wen, Yang Hu, Mingnan Luo 외

Convolution neural netwotks (CNNs) are successfully applied in image recognition task. In this study, we explore the approach of automatic herbal recognition with CNNs and build the standard Chinese herbs datasets firstl…

Deep Sensor Fusion with Pyramid Fusion Networks for 3D Semantic Segmentation

2022-05-26 · Hannah Schieber, Fabian Duerr, Torsten Schoen, Jürgen Beyerer

Robust environment perception for autonomous vehicles is a tremendous challenge, which makes a diverse sensor set with e.g. camera, lidar and radar crucial. In the process of understanding the recorded sensor data, 3D se…

3D Semantic SegmentationAutonomous VehiclesSegmentationSemantic Segmentation+1