paper-with-me

홈 › Papers

CFSum: A Coarse-to-Fine Contribution Network for Multimodal Summarization

2023-07-06 · Min Xiao, Junnan Zhu, Haitao Lin, Yu Zhou, Chengqing Zong

Multimodal summarization usually suffers from the problem that the contribution of the visual modality is unclear. Existing multimodal summarization approaches focus on designing the fusion methods of different modalities, while ignoring the adaptive conditions under which visual modalities are useful. Therefore, we propose a novel Coarse-to-Fine contribution network for multimodal Summarization (CFSum) to consider different contributions of images for summarization. First, to eliminate the interference of useless images, we propose a pre-filter module to abandon useless images. Second, to make accurate use of useful images, we propose two levels of visual complement modules, word level and phrase level. Specifically, image contributions are calculated and are adopted to guide the attention of both textual and visual modalities. Experimental results have shown that CFSum significantly outperforms multiple strong baselines on the standard benchmark. Furthermore, the analysis verifies that useful images can even help generate non-visual words which are implicitly represented in the image.

📄 PDF Abstract BibTeX arXiv:2307.02716

Code (1)

xiaomin418/cfsum 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

CFSum: A Transformer-Based Multi-Modal Video Summarization Framework With Coarse-Fine Fusion

2025-03-01 · Yaowei Guo, Jiazheng Xing, Xiaojun Hou, Shuo Xin 외

Video summarization, by selecting the most informative and/or user-relevant parts of original videos to create concise summary videos, has high research value and consumer demand in today's video proliferation era. Multi…

Video Summarization

Leveraging Entity Information for Cross-Modality Correlation Learning: The Entity-Guided Multimodal Summarization

2024-08-06 · Yanghai Zhang, Ye Liu, Shiwei Wu, Kai Zhang 외

The rapid increase in multimedia data has spurred advancements in Multimodal Summarization with Multimodal Output (MSMO), which aims to produce a multimodal summary that integrates both text and relevant images. The inhe…

Knowledge DistillationLanguage ModelingLanguage Modelling

Progressive Video Summarization via Multimodal Self-supervised Learning

2022-01-07 · Li Haopeng, Ke Qiuhong, Gong Mingming, Tom Drummond

Modern video summarization methods are based on deep neural networks that require a large amount of annotated data for training. However, existing datasets for video summarization are small-scale, easily leading to over-…

Self-Supervised LearningSupervised Video SummarizationVideo ClassificationVideo Summarization

Coarse-to-Fine Attention Models for Document Summarization

2017-09-01 · WS 2017 9 · Jeffrey Ling, Alex Rush, er

Sequence-to-sequence models with attention have been successful for a variety of NLP problems, but their speed does not scale well for tasks with long source sequences such as document summarization. We propose a novel c…

Document SummarizationMachine TranslationQuestion Answering

An Efficient Coarse-to-Fine Facet-Aware Unsupervised Summarization Framework based on Semantic Blocks

2022-08-17 · COLING 2022 10 · Xinnian Liang, Jing Li, Shuangzhi Wu, Jiali Zeng 외

Unsupervised summarization methods have achieved remarkable results by incorporating representations from pre-trained language models. However, existing methods fail to consider efficiency and effectiveness at the same t…

Document Summarization