paper-with-me

홈 › Papers

Which Information Matters? Dissecting Human-written Multi-document Summaries with Partial Information Decomposition

2024-05-23 · Laura Mascarell, Yan L'Homme, Majed El Helou

Understanding the nature of high-quality summaries is crucial to further improve the performance of multi-document summarization. We propose an approach to characterize human-written summaries using partial information decomposition, which decomposes the mutual information provided by all source documents into union, redundancy, synergy, and unique information. Our empirical analysis on different MDS datasets shows that there is a direct dependency between the number of sources and their contribution to the summary.

📄 PDF Abstract BibTeX arXiv:2405.14470

Code (1)

mediatechnologycenter/spider 공식 구현 pytorch

Tasks

Document SummarizationMulti-Document Summarization

Similar Papers 제목 키워드 기반

Towards Tailored Recovery of Lexical Diversity in Literary Machine Translation

2024-08-30 · Esther Ploeger, Huiyuan Lai, Rik van Noord, Antonio Toral

Machine translations are found to be lexically poorer than human translations. The loss of lexical diversity through MT poses an issue in the automatic translation of literature, where it matters not only what is written…

DiversityMachine TranslationRerankingTranslation

The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts

2024-01-23 · Lingfeng Shen, Weiting Tan, Sihao Chen, Yunmo Chen 외

As the influence of large language models (LLMs) spans across global communities, their safety challenges in multilingual settings become paramount for alignment research. This paper examines the variations in safety cha…

Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring

2024-07-04 · Xuansheng Wu, Padmaja Pravin Saraf, Gyeonggeon Lee, Ehsan Latif 외

Large language models (LLMs) have demonstrated strong potential in performing automatic scoring for constructed response assessments. While constructed responses graded by humans are usually based on given grading rubric…

Logical Reasoning

MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts

2024-06-18 · Dominik Macko, Jakub Kopal, Robert Moro, Ivan Srba

Recent LLMs are able to generate high-quality multilingual texts, indistinguishable for humans from authentic human-written ones. Research in machine-generated text detection is however mostly focused on the English lang…

ArticlesBenchmarkingText Detection

Fully Synthetic Data Improves Neural Machine Translation with Knowledge Distillation

2020-12-31 · Alham Fikri Aji, Kenneth Heafield

This paper explores augmenting monolingual data for knowledge distillation in neural machine translation. Source language monolingual text can be incorporated as a forward translation. Interestingly, we find the best way…

Knowledge DistillationMachine TranslationTranslation