Grouping business news stories based on salience of named entities
In news aggregation systems focused on broad news domains, certain stories may appear in multiple articles. Depending on the relative importance of the story, the number of versions can reach dozens or hundreds within a day. The text in these versions may be nearly identical or quite different. Linking multiple versions of a story into a single group brings several important benefits to the end-user{--}reducing the cognitive load on the reader, as well as signaling the relative importance of the story. We present a grouping algorithm, and explore several vector-based representations of input documents: from a baseline using keywords, to a method using salience{--}a measure of importance of named entities in the text. We demonstrate that features beyond keywords yield substantial improvements, verified on a manually-annotated corpus of business news stories.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesWord EmbeddingsSimilar Papers 제목 키워드 기반
WN-Salience: A Corpus of News Articles with Entity Salience Annotations
Entities can be found in various text genres, ranging from tweets and web pages to user queries submitted to web search engines. Existing research either considers all entities in the text equally important, or heuristic…
ArticlesEntity LinkingStory Disambiguation: Tracking Evolving News Stories across News and Social Streams
Following a particular news story online is an important but difficult task, as the relevant information is often scattered across different domains/sources (e.g., news articles, blogs, comments, tweets), presented in va…
ArticlesEntity DisambiguationLearning-To-RankComparing Attitudes to Climate Change in the Media using sentiment analysis based on Latent Dirichlet Allocation
News media typically present biased accounts of news stories, and different publications present different angles on the same event. In this research, we investigate how different publications differ in their approach to…
ArticlesSentiment AnalysisMemory and Knowledge Augmented Language Models for Inferring Salience in Long-Form Stories
Measuring event salience is essential in the understanding of stories. This paper takes a recent unsupervised method for salience detection derived from Barthes Cardinal Functions and theories of surprise and applies it …
FormLanguage ModelingLanguage ModellingRetrieval+1Great Expectations: Unsupervised Inference of Suspense, Surprise and Salience in Storytelling
Stories interest us not because they are a sequence of mundane and predictable events but because they have drama and tension. Crucial to creating dramatic and exciting stories are surprise and suspense. The thesis train…
Deep Learning