paper-with-me

Papers

Analyzing the Dialect Diversity in Multi-document Summaries

2022-10-01 · COLING 2022 10 · Olubusayo Olabisi, Aaron Hudson, Antonie Jetter, Ameeta Agrawal

Social media posts provide a compelling, yet challenging source of data of diverse perspectives from many socially salient groups. Automatic text summarization algorithms make this data accessible at scale by compressing large collections of documents into short summaries that preserve salient information from the source text. In this work, we take a complementary approach to analyzing and improving the quality of summaries generated from social media data in terms of their ability to represent salient as well as diverse perspectives. We introduce a novel dataset, DivSumm, of dialect diverse tweets and human-written extractive and abstractive summaries. Then, we study the extent of dialect diversity reflected in human-written reference summaries as well as system-generated summaries. The results of our extensive experiments suggest that humans annotate fairly well-balanced dialect diverse summaries, and that cluster-based pre-processing approaches seem beneficial in improving the overall quality of the system-generated summaries without loss in diversity.

📄 PDF Abstract BibTeX

Code (1)

portnlp/divsumm 공식 구현

Tasks

DiversityText Summarization

Similar Papers 제목 키워드 기반

Understanding Position Bias Effects on Fairness in Social Multi-Document Summarization

2024-05-03 · Olubusayo Olabisi, Ameeta Agrawal

Text summarization models have typically focused on optimizing aspects of quality such as fluency, relevance, and coherence, particularly in the context of news articles. However, summarization models are increasingly be…

ArticlesDocument SummarizationFairnessMulti-Document Summarization+3

Dialect Diversity in Text Summarization on Twitter

2020-07-15 · Vijay Keswani, L. Elisa Celis

Discussions on Twitter involve participation from different communities with different dialects and it is often necessary to summarize a large number of posts into a representative sample to provide a synopsis. Yet, any …

AttributeDiversityExtractive SummarizationLanguage Identification+1

Unification of Balti and trans-border sister dialects in the essence of LLMs and AI Technology

2024-11-20 · Muhammad Sharif, Jiangyan Yi, Muhammad Shoaib

The language called Balti belongs to the Sino-Tibetan, specifically the Tibeto-Burman language family. It is understood with variations, across populations in India, China, Pakistan, Nepal, Tibet, Burma, and Bhutan, infl…

Diversity

Faithful, Unfaithful or Ambiguous? Multi-Agent Debate with Initial Stance for Summary Evaluation

2025-02-12 · Mahnaz Koupaee, Jake W. Vincent, Saab Mansour, Igor Shalyminov 외

Faithfulness evaluators based on large language models (LLMs) are often fooled by the fluency of the text and struggle with identifying errors in the summaries. We propose an approach to summary faithfulness evaluation i…

Diversity

RoLargeSum: A Large Dialect-Aware Romanian News Dataset for Summary, Headline, and Keyword Generation

2024-12-15 · Andrei-Marius Avram, Mircea Timpuriu, Andreea Iuga, Vlad-Cristian Matei 외

Using supervised automatic summarisation methods requires sufficient corpora that include pairs of documents and their summaries. Similarly to many tasks in natural language processing, most of the datasets available for…

ArticlesBenchmarking