paper-with-me

홈 › Papers

Content Reduction, Surprisal and Information Density Estimation for Long Documents

2023-09-12 · Shaoxiong Ji, Wei Sun, Pekka Marttinen

Many computational linguistic methods have been proposed to study the information content of languages. We consider two interesting research questions: 1) how is information distributed over long documents, and 2) how does content reduction, such as token selection and text summarization, affect the information density in long documents. We present four criteria for information density estimation for long documents, including surprisal, entropy, uniform information density, and lexical density. Among those criteria, the first three adopt the measures from information theory. We propose an attention-based word selection method for clinical notes and study machine summarization for multiple-domain documents. Our findings reveal the systematic difference in information density of long text in various domains. Empirical results on automated medical coding from long clinical notes show the effectiveness of the attention-based word selection method.

📄 PDF Abstract BibTeX arXiv:2309.06009

Code (0)

등록된 구현이 없습니다.

Tasks

Density EstimationText Summarization

Similar Papers 제목 키워드 기반

Controlling Surprisal in Music Generation via Information Content Curve Matching

2024-08-12 · Mathias Rose Bjare, Stefan Lattner, Gerhard Widmer

In recent years, the quality and public interest in music generation systems have grown, encouraging research into various ways to control these systems. We propose a novel method for controlling surprisal in music gener…

Music Generation

Is Information Density Uniform when Utterances are Grounded on Perception and Discourse?

2026-02-16 · Matteo Gay, Coleman Haley, Mario Giulianelli, Edoardo Ponti arxiv

The Uniform Information Density (UID) hypothesis posits that speakers are subject to a communicative pressure to distribute information evenly within utterances, minimising surprisal variance. While this hypothesis has b…

Visual Storytelling

Expect the Unexpected? Testing the Surprisal of Salient Entities

2026-04-12 · Jessica Lin, Amir Zeldes arxiv

Previous work examining the Uniform Information Density (UID) hypothesis has shown that while information as measured by surprisal metrics is distributed more or less evenly across documents overall, local discrepancies …

Information Density as a Factor for Variation in the Embedding of Relative Clauses

2017-05-18 · Augustin Speyer, Robin Lemke

In German, relative clauses can be positioned in-situ or extraposed. A potential factor for the variation might be information density. In this study, this hypothesis is tested with a corpus of 17th century German funera…

Language ModelingLanguage Modelling

How is BERT surprised? Layerwise detection of linguistic anomalies

2021-05-16 · ACL 2021 5 · Bai Li, Zining Zhu, Guillaume Thomas, Yang Xu 외

Transformer language models have shown remarkable ability in detecting when a word is anomalous in context, but likelihood scores offer no information about the cause of the anomaly. In this work, we use Gaussian models …

Density Estimation