French Coreference for Spoken and Written Language
Coreference resolution aims at identifying and grouping all mentions referring to the same entity. In French, most systems run different setups, making their comparison difficult. In this paper, we present an extensive comparison of several coreference resolution systems for French. The systems have been trained on two corpora (ANCOR for spoken language and Democrat for written language) annotated with coreference chains, and augmented with syntactic and semantic information. The models are compared with different configurations (e.g. with and without singletons). In addition, we evaluate mention detection and coreference resolution apart. We present a full-stack model that outperforms other approaches. This model allows us to study the impact of mention detection errors on coreference resolution. Our analysis shows that mention detection can be improved by focusing on boundary identification while advances in the pronoun-noun relation detection can help the coreference task. Another contribution of this work is the first end-to-end neural French coreference resolution model trained on Democrat (written texts), which compares to the state-of-the-art systems for oral French.
Code (0)
등록된 구현이 없습니다.
Tasks
coreference-resolutionCoreference ResolutionSimilar Papers 제목 키워드 기반
ANCOR\_Centre, a large free spoken French coreference corpus: description of the resource and reliability measures
This article presents ANCOR{\_}Centre, a French coreference corpus, available under the Creative Commons Licence. With a size of around 500,000 words, the corpus is large enough to serve the needs of data-driven approach…
Coreference ResolutionEntity LinkingInformation RetrievalVariation in Coreference Strategies across Genres and Production Media
In response to (i) inconclusive results in the literature as to the properties of coreference chains in written versus spoken language, and (ii) a general lack of work on automatic coreference resolution on both spoken l…
coreference-resolutionCoreference ResolutionCoreference in Spoken vs. Written Texts: a Corpus-based Analysis
This paper describes an empirical study of coreference in spoken vs. written text. We focus on the comparison of two particular text types, interviews and popular science texts, as instances of spoken and written texts s…
coreference-resolutionCoreference ResolutionDiversityAssessing Human Translations from French to Bambara for Machine Learning: a Pilot Study
We present novel methods for assessing the quality of human-translated aligned texts for learning machine translation models of under-resourced languages. Malian university students translated French texts, producing eit…
BIG-bench Machine LearningMachine TranslationTranslationExploring Coreference Features in Heterogeneous Data
The present paper focuses on variation phenomena in coreference chains. We address the hypothesis that the degree of structural variation between chain elements depends on language-specific constraints and preferences an…
coreference-resolutionCoreference Resolutionfeature selectionMachine Translation+3