paper-with-me

홈 › Papers

Earlier Isn't Always Better: Sub-aspect Analysis on Corpus and System Biases in Summarization

2019-08-30 · IJCNLP 2019 11 · Taehee Jung, Dongyeop Kang, Lucas Mentch, Eduard Hovy

Despite the recent developments on neural summarization systems, the underlying logic behind the improvements from the systems and its corpus-dependency remains largely unexplored. Position of sentences in the original text, for example, is a well known bias for news summarization. Following in the spirit of the claim that summarization is a combination of sub-functions, we define three sub-aspects of summarization: position, importance, and diversity and conduct an extensive analysis of the biases of each sub-aspect with respect to the domain of nine different summarization corpora (e.g., news, academic papers, meeting minutes, movie script, books, posts). We find that while position exhibits substantial bias in news articles, this is not the case, for example, with academic papers and meeting minutes. Furthermore, our empirical study shows that different types of summarization systems (e.g., neural-based) are composed of different degrees of the sub-aspects. Our study provides useful lessons regarding consideration of underlying sub-aspects when collecting a new summarization dataset or developing a new system.

📄 PDF Abstract BibTeX arXiv:1908.11723

Code (1)

dykang/biassum pytorch

Tasks

ArticlesDiversityNews SummarizationPosition

Similar Papers 제목 키워드 기반

Enhancing Aspect Extraction for Hindi

2021-08-01 · ACL (ECNLP) 2021 8 · Arghya Bhattacharya, Alok Debnath, Manish Shrivastava

Aspect extraction is not a well-explored topic in Hindi, with only one corpus having been developed for the task. In this paper, we discuss the merits of the existing corpus in terms of quality, size, sparsity, and perfo…

Aspect-Based Sentiment AnalysisAspect-Based Sentiment Analysis (ABSA)Aspect ExtractionSentiment Analysis

A Multimodal Corpus of Expert Gaze and Behavior during Phonetic Segmentation Tasks

2017-12-13 · LREC 2018 5 · Arif Khan, Ingmar Steiner, Yusuke Sugano, Andreas Bulling 외

Phonetic segmentation is the process of splitting speech into distinct phonetic units. Human experts routinely perform this task manually by analyzing auditory and visual cues using analysis software, which is an extreme…

Segmentation

The Royal Society Corpus 6.0: Providing 300+ Years of Scientific Writing for Humanistic Study

2020-05-01 · LREC 2020 5 · Stefan Fischer, J{\"o}rg Knappen, Katrin Menzel, Elke Teich

We present a new, extended version of the Royal Society Corpus (RSC), a diachronic corpus of scientific English now covering 300+ years of scientific writing (1665--1996). The corpus comprises 47 837 texts, primarily sci…

Articles

UWB at SemEval-2020 Task 1: Lexical Semantic Change Detection

2020-11-30 · SEMEVAL 2020 · Ondřej Pražák, Pavel Přibáň, Stephen Taylor, Jakub Sido

In this paper, we describe our method for the detection of lexical semantic change, i.e., word sense changes over time. We examine semantic differences between specific words in two corpora, chosen from different time pe…

Change DetectionTask 2

Words with Consistent Diachronic Usage Patterns are Learned Earlier: A Computational Analysis Using Temporally Aligned Word Embeddings

2021-04-20 · Cognitive Science 2021 4 · Giovanni Cassani, Federico Bianchi, Marco Marelli

In this study, we use temporally aligned word embeddings and a large diachronic corpus of English to quantify language change in a data-driven, scalable way, which is grounded in language use. We show a unique and reliab…

Diachronic Word EmbeddingsDiversityRelationWord Embeddings