DAWT: Densely Annotated Wikipedia Texts across multiple languages
In this work, we open up the DAWT dataset - Densely Annotated Wikipedia Texts across multiple languages. The annotations include labeled text mentions mapping to entities (represented by their Freebase machine ids) as well as the type of the entity. The data set contains total of 13.6M articles, 5.0B tokens, 13.8M mention entity co-occurrences. DAWT contains 4.8 times more anchor text to entity links than originally present in the Wikipedia markup. Moreover, it spans several languages including English, Spanish, Italian, German, French and Arabic. We also present the methodology used to generate the dataset which enriches Wikipedia markup in order to increase number of links. In addition to the main dataset, we open up several derived datasets including mention entity co-occurrence counts and entity embeddings, as well as mappings between Freebase ids and Wikidata item ids. We also discuss two applications of these datasets and hope that opening them up would prove useful for the Natural Language Processing and Information Retrieval communities, as well as facilitate multi-lingual research.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesEntity EmbeddingsInformation RetrievalRetrievalSimilar Papers 제목 키워드 기반
WikiCoref: An English Coreference-annotated Corpus of Wikipedia Articles
This paper presents WikiCoref, an English corpus annotated for anaphoric relations, where all documents are from the English version of Wikipedia. Our annotation scheme follows the one of OntoNotes with a few disparities…
Articlescoreference-resolutionCoreference ResolutionDBpedia Abstracts: A Large-Scale, Open, Multilingual NLP Training Corpus
The ever increasing importance of machine learning in Natural Language Processing is accompanied by an equally increasing need in large-scale training and evaluation corpora. Due to its size, its openness and relative qu…
Entity LinkingMultilingual NLPRecognizing Descriptive Wikipedia Categories for Historical Figures
Wikipedia is a useful knowledge source that benefits many applications in language processing and knowledge representation. An important feature of Wikipedia is that of categories. Wikipedia pages are assigned different …
DescriptiveInformation RetrievalRetrievalTAGDDisCo: A Discourse Coherence Dataset for Danish
To date, there has been no resource for studying discourse coherence on real-world Danish texts. Discourse coherence has mostly been approached with the assumption that incoherent texts can be represented by coherent tex…
pioNER: Datasets and Baselines for Armenian Named Entity Recognition
In this work, we tackle the problem of Armenian named entity recognition, providing silver- and gold-standard datasets as well as establishing baseline results on popular models. We present a 163000-token named entity co…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Word Embeddings