Content Reduction, Surprisal and Information Density Estimation for Long Documents
Many computational linguistic methods have been proposed to study the information content of languages. We consider two interesting research questions: 1) how is information distributed over long documents, and 2) how does content reduction, such as token selection and text summarization, affect the information density in long documents. We present four criteria for information density estimation for long documents, including surprisal, entropy, uniform information density, and lexical density. Among those criteria, the first three adopt the measures from information theory. We propose an attention-based word selection method for clinical notes and study machine summarization for multiple-domain documents. Our findings reveal the systematic difference in information density of long text in various domains. Empirical results on automated medical coding from long clinical notes show the effectiveness of the attention-based word selection method.
Code (0)
등록된 구현이 없습니다.
Tasks
Density EstimationText SummarizationSimilar Papers 제목 키워드 기반
Controlling Surprisal in Music Generation via Information Content Curve Matching
In recent years, the quality and public interest in music generation systems have grown, encouraging research into various ways to control these systems. We propose a novel method for controlling surprisal in music gener…
Music GenerationIs Information Density Uniform when Utterances are Grounded on Perception and Discourse?
The Uniform Information Density (UID) hypothesis posits that speakers are subject to a communicative pressure to distribute information evenly within utterances, minimising surprisal variance. While this hypothesis has b…
Visual StorytellingExpect the Unexpected? Testing the Surprisal of Salient Entities
Previous work examining the Uniform Information Density (UID) hypothesis has shown that while information as measured by surprisal metrics is distributed more or less evenly across documents overall, local discrepancies …
Information Density as a Factor for Variation in the Embedding of Relative Clauses
In German, relative clauses can be positioned in-situ or extraposed. A potential factor for the variation might be information density. In this study, this hypothesis is tested with a corpus of 17th century German funera…
Language ModelingLanguage ModellingHow is BERT surprised? Layerwise detection of linguistic anomalies
Transformer language models have shown remarkable ability in detecting when a word is anomalous in context, but likelihood scores offer no information about the cause of the anomaly. In this work, we use Gaussian models …
Density Estimation