paper-with-me

홈 › Papers

GUM-SAGE: A Novel Dataset and Approach for Graded Entity Salience Prediction

2025-04-15 · Jessica Lin, Amir Zeldes

Determining and ranking the most salient entities in a text is critical for user-facing systems, especially as users increasingly rely on models to interpret long documents they only partially read. Graded entity salience addresses this need by assigning entities scores that reflect their relative importance in a text. Existing approaches fall into two main categories: subjective judgments of salience, which allow for gradient scoring but lack consistency, and summarization-based methods, which define salience as mention-worthiness in a summary, promoting explainability but limiting outputs to binary labels (entities are either summary-worthy or not). In this paper, we introduce a novel approach for graded entity salience that combines the strengths of both approaches. Using an English dataset spanning 12 spoken and written genres, we collect 5 summaries per document and calculate each entity's salience score based on its presence across these summaries. Our approach shows stronger correlation with scores based on human summaries and alignments, and outperforms existing techniques, including LLMs. We release our data and code at https://github.com/jl908069/gum_sum_salience to support further research on graded salient entity extraction.

📄 PDF Abstract BibTeX arXiv:2504.10792

Code (1)

jl908069/gum_sum_salience 공식 구현

Similar Papers 제목 키워드 기반

Not Worth Mentioning? A Pilot Study on Salient Proposition Annotation

2026-03-28 · Amir Zeldes, Katherine Conhaim, Lauren Levine arxiv

Despite a long tradition of work on extractive summarization, which by nature aims to recover the most important propositions in a text, little work has been done on operationalizing graded proposition salience in natura…

Discourse Parsing

What makes an entity salient in discourse?

2025-08-22 · Amir Zeldes, Jessica Lin arxiv

Entities in discourse vary in salience: main participants, objects and locations stay prominent, while others are quickly forgotten, raising questions about how humans signal and infer discourse-level salience. Using a g…

WN-Salience: A Corpus of News Articles with Entity Salience Annotations

2020-05-01 · LREC 2020 5 · Chuan Wu, Evangelos Kanoulas, Maarten de Rijke, Wei Lu

Entities can be found in various text genres, ranging from tweets and web pages to user queries submitted to web search engines. Existing research either considers all entities in the text equally important, or heuristic…

ArticlesEntity Linking

Towards Better Text Understanding and Retrieval through Kernel Entity Salience Modeling

2018-05-03 · Chenyan Xiong, Zhengzhong Liu, Jamie Callan, Tie-Yan Liu

This paper presents a Kernel Entity Salience Model (KESM) that improves text understanding and retrieval by better estimating entity salience (importance) in documents. KESM represents entities by knowledge enriched dist…

Retrieval

Crowdsourced Corpus with Entity Salience Annotations

2016-05-01 · LREC 2016 5 · Milan Dojchinovski, Dinesh Reddy, Tom{\'a}{\v{s}} Kliegr, Tom{\'a}{\v{s}} Vitvar 외

In this paper, we present a crowdsourced dataset which adds entity salience (importance) annotations to the Reuters-128 dataset, which is subset of Reuters-21578. The dataset is distributed under a free license and publi…