paper-with-me

Papers

Semantics-Consistent Cross-domain Summarization via Optimal Transport Alignment

2022-10-10 · JieLin Qiu, Jiacheng Zhu, Mengdi Xu, Franck Dernoncourt, Trung Bui, Zhaowen Wang, Bo Li, Ding Zhao, Hailin Jin

Multimedia summarization with multimodal output (MSMO) is a recently explored application in language grounding. It plays an essential role in real-world applications, i.e., automatically generating cover images and titles for news articles or providing introductions to online videos. However, existing methods extract features from the whole video and article and use fusion methods to select the representative one, thus usually ignoring the critical structure and varying semantics. In this work, we propose a Semantics-Consistent Cross-domain Summarization (SCCS) model based on optimal transport alignment with visual and textual segmentation. In specific, our method first decomposes both video and article into segments in order to capture the structural semantics, respectively. Then SCCS follows a cross-domain alignment objective with optimal transport distance, which leverages multimodal interaction to match and select the visual and textual summary. We evaluated our method on three recent multimodal datasets and demonstrated the effectiveness of our method in producing high-quality multimodal summaries.

📄 PDF Abstract BibTeX arXiv:2210.04722

Code (0)

등록된 구현이 없습니다.

Tasks

Articlesmultimodal interaction

Similar Papers 제목 키워드 기반

Efficient Extractive Summarization with MAMBA-Transformer Hybrids for Low-Resource Scenarios

2026-03-01 · Nisrine Ait Khayi arxiv

Extractive summarization of long documents is bottlenecked by quadratic complexity, often forcing truncation and limiting deployment in resource-constrained settings. We introduce the first Mamba-Transformer hybrid for e…

DWTSumm: Discrete Wavelet Transform for Document Summarization

2026-04-22 · Rana Salama, Abdou Youssef, Mona Diab arxiv

Summarizing long, domain-specific documents with large language models (LLMs) remains challenging due to context limitations, information loss, and hallucinations, particularly in clinical and legal settings. We propose …

Document SummarizationSemantic Similarity

Cut to the Chase: Training-free Multimodal Summarization via Chain-of-Events

2026-03-06 · Xiaoxing You, Qiang Huang, Lingyu Li, Xiaojun Chang 외 arxiv

Multimodal Summarization (MMS) aims to generate concise textual summaries by understanding and integrating information across videos, transcripts, and images. However, existing approaches still suffer from three main cha…

Domain Generalization

MHMS: Multimodal Hierarchical Multimedia Summarization

2022-04-07 · JieLin Qiu, Jiacheng Zhu, Mengdi Xu, Franck Dernoncourt 외

Multimedia summarization with multimodal output can play an essential role in real-world applications, i.e., automatically generating cover images and titles for news articles or providing introductions to online videos.…

Articles

Enriching and Controlling Global Semantics for Text Summarization

2021-09-22 · EMNLP 2021 11 · Thong Nguyen, Anh Tuan Luu, Truc Lu, Tho Quan

Recently, Transformer-based models have been proven effective in the abstractive summarization task by creating fluent and informative summaries. Nevertheless, these models still suffer from the short-range dependency pr…

Abstractive Text SummarizationText GenerationText Summarization