paper-with-me

Papers

MHMS: Multimodal Hierarchical Multimedia Summarization

2022-04-07 · JieLin Qiu, Jiacheng Zhu, Mengdi Xu, Franck Dernoncourt, Trung Bui, Zhaowen Wang, Bo Li, Ding Zhao, Hailin Jin

Multimedia summarization with multimodal output can play an essential role in real-world applications, i.e., automatically generating cover images and titles for news articles or providing introductions to online videos. In this work, we propose a multimodal hierarchical multimedia summarization (MHMS) framework by interacting visual and language domains to generate both video and textual summaries. Our MHMS method contains video and textual segmentation and summarization module, respectively. It formulates a cross-domain alignment objective with optimal transport distance which leverages cross-domain interaction to generate the representative keyframe and textual summary. We evaluated MHMS on three recent multimodal datasets and demonstrated the effectiveness of our method in producing high-quality multimodal summaries.

📄 PDF Abstract BibTeX arXiv:2204.03734

Code (0)

등록된 구현이 없습니다.

Tasks

Articles

Similar Papers 제목 키워드 기반

MSMO: Multimodal Summarization with Multimodal Output

2018-10-01 · EMNLP 2018 10 · Junnan Zhu, Haoran Li, Tianshang Liu, Yu Zhou 외

Multimodal summarization has drawn much attention due to the rapid growth of multimedia data. The output of the current multimodal summarization systems is usually represented in texts. However, we have found through exp…

InformativenessText Summarization

MSCMHMST: A traffic flow prediction model based on Transformer

2025-03-16 · Weiyang Geng, Yiming Pan, Zhecong Xing, Dongyu Liu 외

This study proposes a hybrid model based on Transformers, named MSCMHMST, aimed at addressing key challenges in traffic flow prediction. Traditional single-method approaches show limitations in traffic prediction tasks, …

PredictionTraffic Prediction

UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation

2021-09-13 · Zhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang 외

With the rapid increase of multimedia data, a large body of literature has emerged to work on multimodal summarization, the majority of which target at refining salient information from textual and visual modalities to o…

Abstractive Text SummarizationDecoderImage CaptioningKnowledge Distillation+1

Leveraging Entity Information for Cross-Modality Correlation Learning: The Entity-Guided Multimodal Summarization

2024-08-06 · Yanghai Zhang, Ye Liu, Shiwei Wu, Kai Zhang 외

The rapid increase in multimedia data has spurred advancements in Multimodal Summarization with Multimodal Output (MSMO), which aims to produce a multimodal summary that integrates both text and relevant images. The inhe…

Knowledge DistillationLanguage ModelingLanguage Modelling

Hierarchical3D Adapters for Long Video-to-text Summarization

2022-10-10 · Pinelopi Papalampidi, Mirella Lapata

In this paper, we focus on video-to-text summarization and investigate how to best utilize multimodal information for summarizing long inputs (e.g., an hour-long TV show) into long outputs (e.g., a multi-sentence summary…

SentenceText Summarization