paper-with-me

홈 › Papers

Unifying Cross-lingual Summarization and Machine Translation with Compression Rate

2021-10-15 · Yu Bai, Heyan Huang, Kai Fan, Yang Gao, Yiming Zhu, Jiaao Zhan, Zewen Chi, Boxing Chen

Cross-Lingual Summarization (CLS) is a task that extracts important information from a source document and summarizes it into a summary in another language. It is a challenging task that requires a system to understand, summarize, and translate at the same time, making it highly related to Monolingual Summarization (MS) and Machine Translation (MT). In practice, the training resources for Machine Translation are far more than that for cross-lingual and monolingual summarization. Thus incorporating the Machine Translation corpus into CLS would be beneficial for its performance. However, the present work only leverages a simple multi-task framework to bring Machine Translation in, lacking deeper exploration. In this paper, we propose a novel task, Cross-lingual Summarization with Compression rate (CSC), to benefit Cross-Lingual Summarization by large-scale Machine Translation corpus. Through introducing compression rate, the information ratio between the source and the target text, we regard the MT task as a special CLS task with a compression rate of 100%. Hence they can be trained as a unified task, sharing knowledge more effectively. However, a huge gap exists between the MT task and the CLS task, where samples with compression rates between 30% and 90% are extremely rare. Hence, to bridge these two tasks smoothly, we propose an effective data augmentation method to produce document-summary pairs with different compression rates. The proposed method not only improves the performance of the CLS task, but also provides controllability to generate summaries in desired lengths. Experiments demonstrate that our method outperforms various strong baselines in three cross-lingual summarization datasets. We released our code and data at https://github.com/ybai-nlp/CLS_CR.

📄 PDF Abstract BibTeX arXiv:2110.07936

Code (1)

ybai-nlp/cls_cr 공식 구현 pytorch

Tasks

Data AugmentationMachine TranslationTranslation

Similar Papers 제목 키워드 기반

A Deep Reinforced Model for Zero-Shot Cross-Lingual Summarization with Bilingual Semantic Similarity Rewards

2020-06-27 · WS 2020 7 · Zi-Yi Dou, Sachin Kumar, Yulia Tsvetkov

Cross-lingual text summarization aims at generating a document summary in one language given input in another language. It is a practically important but under-explored task, primarily due to the dearth of available data…

Machine Translationreinforcement-learningReinforcement LearningReinforcement Learning (RL)+4

Multi-Task Learning for Cross-Lingual Abstractive Summarization

2020-10-15 · LREC 2022 6 · Sho Takase, Naoaki Okazaki

We present a multi-task learning framework for cross-lingual abstractive summarization to augment training data. Recent studies constructed pseudo cross-lingual abstractive summarization data to train their neural encode…

Abstractive Text SummarizationCross-Lingual Abstractive SummarizationMachine TranslationMulti-Task Learning+2

NCLS: Neural Cross-Lingual Summarization

2019-08-31 · IJCNLP 2019 11 · Junnan Zhu, Qian Wang, Yining Wang, Yu Zhou 외

Cross-lingual summarization (CLS) is the task to produce a summary in one particular language for a source document in a different language. Existing methods simply divide this task into two steps: summarization and tran…

Machine TranslationMulti-Task LearningTranslation

WikiLingua: A New Benchmark Dataset for Cross-Lingual Abstractive Summarization

2020-10-07 · Findings of the Association for Computational Linguistics 2020 · Faisal Ladhak, Esin Durmus, Claire Cardie, Kathleen McKeown

We introduce WikiLingua, a large-scale, multilingual dataset for the evaluation of crosslingual abstractive summarization systems. We extract article and summary pairs in 18 languages from WikiHow, a high quality, collab…

Abstractive Text SummarizationCross-Lingual Abstractive SummarizationMachine TranslationTranslation

MT6: Multilingual Pretrained Text-to-Text Transformer with Translation Pairs

2021-04-18 · EMNLP 2021 11 · Zewen Chi, Li Dong, Shuming Ma, Shaohan Huang Xian-Ling Mao 외

Multilingual T5 (mT5) pretrains a sequence-to-sequence model on massive monolingual texts, which has shown promising results on many cross-lingual tasks. In this paper, we improve multilingual text-to-text transfer Trans…

Abstractive Text SummarizationMachine Translationnamed-entity-recognitionNamed Entity Recognition+5