paper-with-me

Papers

Knowledge-enriched Attention Network with Group-wise Semantic for Visual Storytelling

2022-03-10 · Tengpeng Li, Hanli Wang, Bin He, Chang Wen Chen

As a technically challenging topic, visual storytelling aims at generating an imaginary and coherent story with narrative multi-sentences from a group of relevant images. Existing methods often generate direct and rigid descriptions of apparent image-based contents, because they are not capable of exploring implicit information beyond images. Hence, these schemes could not capture consistent dependencies from holistic representation, impairing the generation of reasonable and fluent story. To address these problems, a novel knowledge-enriched attention network with group-wise semantic model is proposed. Three main novel components are designed and supported by substantial experiments to reveal practical advantages. First, a knowledge-enriched attention network is designed to extract implicit concepts from external knowledge system, and these concepts are followed by a cascade cross-modal attention mechanism to characterize imaginative and concrete representations. Second, a group-wise semantic module with second-order pooling is developed to explore the globally consistent guidance. Third, a unified one-stage story generation model with encoder-decoder structure is proposed to simultaneously train and infer the knowledge-enriched attention network, group-wise semantic module and multi-modal story generation decoder in an end-to-end fashion. Substantial experiments on the popular Visual Storytelling dataset with both objective and subjective evaluation metrics demonstrate the superior performance of the proposed scheme as compared with other state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2203.05346

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderStory GenerationVisual Storytelling

Similar Papers 제목 키워드 기반

Harnessing Textual Semantic Priors for Knowledge Transfer and Refinement in CLIP-Driven Continual Learning

2025-08-03 · Lingfeng He, De Cheng, Di Xu, Huaijie Wang 외 arxiv

Continual learning (CL) aims to equip models with the ability to learn from a stream of tasks without forgetting previous knowledge. With the progress of vision-language models like Contrastive Language-Image Pre-trainin…

Continual Learning

VGSG: Vision-Guided Semantic-Group Network for Text-based Person Search

2023-11-13 · Shuting He, Hao Luo, Wei Jiang, Xudong Jiang 외

Text-based Person Search (TBPS) aims to retrieve images of target pedestrian indicated by textual descriptions. It is essential for TBPS to extract fine-grained local features and align them crossing modality. Existing m…

Person SearchText based Person RetrievalText based Person SearchTransfer Learning

Exploiting Local and Global Structure for Point Cloud Semantic Segmentation with Contextual Point Representations

2019-11-13 · NeurIPS 2019 12 · Xu Wang, Jingming He, Lin Ma

In this paper, we propose one novel model for point cloud semantic segmentation, which exploits both the local and global structures within the point cloud based on the contextual point representations. Specifically, we …

Graph AttentionSemantic Segmentation

Pair-wise Layer Attention with Spatial Masking for Video Prediction

2023-11-19 · Ping Li, Chenhan Zhang, Zheng Yang, Xianghua Xu 외

Video prediction yields future frames by employing the historical frames and has exhibited its great potential in many applications, e.g., meteorological prediction, and autonomous driving. Previous works often decode th…

Autonomous DrivingDecoderPredictionVideo Prediction

Group-Wise Semantic Mining for Weakly Supervised Semantic Segmentation

2020-12-09 · Xueyi Li, Tianfei Zhou, Jianwu Li, Yi Zhou 외

Acquiring sufficient ground-truth supervision to train deep visual models has been a bottleneck over the years due to the data-hungry nature of deep learning. This is exacerbated in some structured prediction tasks, such…

Graph Neural NetworkSegmentationSemantic SegmentationStructured Prediction+2