TESA: A Task in Entity Semantic Aggregation for Abstractive Summarization
Human-written texts contain frequent generalizations and semantic aggregation of content. In a document, they may refer to a pair of named entities such as {}London{'} and {}Paris{'} with different expressions: {`}the major cities{''}, {}the capital cities{''} and {`}two European cities{''}. Yet generation, especially, abstractive summarization systems have so far focused heavily on paraphrasing and simplifying the source content, to the exclusion of such semantic abstraction capabilities. In this paper, we present a new dataset and task aimed at the semantic aggregation of entities. TESA contains a dataset of 5.3K crowd-sourced entity aggregations of Person, Organization, and Location named entities. The aggregations are document-appropriate, meaning that they are produced by annotators to match the situational context of a given news article from the New York Times. We then build baseline models for generating aggregations given a tuple of entities and document context. We finetune on TESA an encoder-decoder language model and compare it with simpler classification methods based on linguistically informed features. Our quantitative and qualitative evaluations show reasonable performance in making a choice from a given list of expressions, but free-form expressions are understandably harder to generate and evaluate.
Code (0)
등록된 구현이 없습니다.
Tasks
Abstractive Text SummarizationDecoderLanguage ModelingLanguage ModellingSimilar Papers 제목 키워드 기반
Source-summary Entity Aggregation in Abstractive Summarization
In a text, entities mentioned earlier can be referred to in later discourse by a more general description. For example, Celine Dion and Justin Bieber can be referred to by Canadian singers or celebrities. In this work, w…
Abstractive Text SummarizationRoute Sparse Autoencoder to Interpret Large Language Models
Mechanistic interpretability of large language models (LLMs) aims to uncover the internal processes of information propagation and reasoning. Sparse autoencoders (SAEs) have demonstrated promise in this domain by extract…
RemoteSAM: Towards Segment Anything for Earth Observation
We aim to develop a robust yet flexible visual foundation model for Earth observation. It should possess strong capabilities in recognizing and localizing diverse visual targets while providing compatibility with various…
AttributeEarth ObservationReferring ExpressionReferring Expression SegmentationEntity-driven Fact-aware Abstractive Summarization of Biomedical Literature
As part of the large number of scientific articles being published every year, the publication rate of biomedical literature has been increasing. Consequently, there has been considerable effort to harness and summarize …
Abstractive Text SummarizationArticlesDecoderDocument Summarization+1CommentsRadar: Dive into Unique Data on All Comments on the Web
We introduce an entity-centric search engineCommentsRadarthatpairs entity queries with articles and user opinions covering a widerange of topics from top commented sites. The engine aggregatesarticles and comments for th…
AllArticles