Extended Named Entities Annotation on OCRed Documents: From Corpus Constitution to Evaluation Campaign
Within the framework of the Quaero project, we proposed a new definition of named entities, based upon an extension of the coverage of named entities as well as the structure of those named entities. In this new definition, the extended named entities we proposed are both hierarchical and compositional. In this paper, we focused on the annotation of a corpus composed of press archives, OCRed from French newspapers of December 1890. We present the methodology we used to produce the corpus and the characteristics of the corpus in terms of named entities annotation. This annotated corpus has been used in an evaluation campaign. We present this evaluation, the metrics we used and the results obtained by the participants.
Code (0)
등록된 구현이 없습니다.
Tasks
Named Entity Recognition (NER)Optical Character Recognition (OCR)Similar Papers 제목 키워드 기반
On the Robustness of Document-Level Relation Extraction Models to Entity Name Variations
Driven by the demand for cross-sentence and large-scale relation extraction, document-level relation extraction (DocRE) has attracted increasing research interest. Despite the continuous improvement in performance, we fi…
Document-level Relation ExtractionIn-Context LearningRelationRelation Extraction+1Old Content and Modern Tools - Searching Named Entities in a Finnish OCRed Historical Newspaper Collection 1771-1910
Named Entity Recognition (NER), search, classification and tagging of names and name like frequent informational elements in texts, has become a standard information extraction procedure for textual data. NER has been ap…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1Revisiting DocRED -- Addressing the False Negative Problem in Relation Extraction
The DocRED dataset is one of the most popular and widely used benchmarks for document-level relation extraction (RE). It adopts a recommend-revise annotation scheme so as to have a large-scale annotated dataset. However,…
Document-level Relation ExtractionRelationRelation ExtractionDocRED: A Large-Scale Document-Level Relation Extraction Dataset
Multiple entities in a document generally exhibit complex inter-sentence relations, and cannot be well handled by existing relation extraction (RE) methods that typically focus on extracting intra-sentence relations for …
Document-level Relation ExtractionRelationRelation ExtractionSentenceDoes Recommend-Revise Produce Reliable Annotations? An Analysis on Missing Instances in DocRED
DocRED is a widely used dataset for document-level relation extraction. In the large-scale annotation, a \textit{recommend-revise} scheme is adopted to reduce the workload. Within this scheme, annotators are provided wit…
Document-level Relation ExtractionRelationRelation Extraction