paper-with-me

Papers

Diverse Word Choices, Same Reference: Annotating Lexically-Rich Cross-Document Coreference

2026-02-19 · Anastasia Zhukova, Felix Hamborg, Karsten Donnay, Norman Meuschke, Bela Gipp arxiv

Cross-document coreference resolution (CDCR) identifies and links mentions of the same entities and events across related documents, enabling content analysis that aggregates information at the level of discourse participants. However, existing datasets primarily focus on event resolution and employ a narrow definition of coreference, which limits their effectiveness in analyzing diverse and polarized news coverage where wording varies widely. This paper proposes a revised CDCR annotation scheme of the NewsWCL50 dataset, treating coreference chains as discourse elements (DEs) and conceptual units of analysis. The approach accommodates both identity and near-identity relations, e.g., by linking "the caravan" - "asylum seekers" - "those contemplating illegal entry", allowing models to capture lexical diversity and framing variation in media discourse, while maintaining the fine-grained annotation of DEs. We reannotate the NewsWCL50 and a subset of ECB+ using a unified codebook and evaluate the new datasets through lexical diversity metrics and a same-head-lemma baseline. The results show that the reannotated datasets align closely, falling between the original ECB+ and NewsWCL50, thereby supporting balanced and discourse-aware CDCR research in the news domain.

📄 PDF Abstract BibTeX arXiv:2602.17424

Code (0)

등록된 구현이 없습니다.

Tasks

Coreference Resolution

Similar Papers 제목 키워드 기반

Code Book for the Annotation of Diverse Cross-Document Coreference of Entities in News Articles

2023-10-18 · Jakob Vogel

This paper presents a scheme for annotating coreference across news articles, extending beyond traditional identity relations by also considering near-identity and bridging relations. It includes a precise description of…

Articles

Annotation-Efficient Preference Optimization for Language Model Alignment

2024-05-22 · Yuu Jinnai, Ukyo Honda

Preference optimization is a standard approach to fine-tuning large language models to align with human preferences. The quality, diversity, and quantity of the preference dataset are critical to the effectiveness of pre…

DiversityLanguage ModelingLanguage Modellingmodel

Annotating the MASC Corpus with BabelNet

2014-05-01 · LREC 2014 5 · Andrea Moro, Roberto Navigli, Francesco Maria Tucci, Rebecca J. Passonneau

In this paper we tackle the problem of automatically annotating, with both word senses and named entities, the MASC 3.0 corpus, a large English corpus covering a wide range of genres of written and spoken text. We use Ba…

Entity LinkingReading ComprehensionRelation ExtractionWord Sense Disambiguation

Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models

2026-08-15 · Varvara Arzt, Allan Hanbury, Terra Blevins arxiv

We systematically compare word order preferences in decoder-only language models across 192 artificial languages and typologically diverse natural languages. On artificial languages, models exhibit a left-branching prefe…

Learning the Preferences of Ignorant, Inconsistent Agents

2015-12-18 · Owain Evans, Andreas Stuhlmueller, Noah D. Goodman

An important use of machine learning is to learn what people value. What posts or photos should a user be shown? Which jobs or activities would a person find rewarding? In each case, observations of people's past choices…

Decision Making