paper-with-me

홈 › Papers

DAWT: Densely Annotated Wikipedia Texts across multiple languages

2017-03-02 · Nemanja Spasojevic, Preeti Bhargava, Guoning Hu

In this work, we open up the DAWT dataset - Densely Annotated Wikipedia Texts across multiple languages. The annotations include labeled text mentions mapping to entities (represented by their Freebase machine ids) as well as the type of the entity. The data set contains total of 13.6M articles, 5.0B tokens, 13.8M mention entity co-occurrences. DAWT contains 4.8 times more anchor text to entity links than originally present in the Wikipedia markup. Moreover, it spans several languages including English, Spanish, Italian, German, French and Arabic. We also present the methodology used to generate the dataset which enriches Wikipedia markup in order to increase number of links. In addition to the main dataset, we open up several derived datasets including mention entity co-occurrence counts and entity embeddings, as well as mappings between Freebase ids and Wikidata item ids. We also discuss two applications of these datasets and hope that opening them up would prove useful for the Natural Language Processing and Information Retrieval communities, as well as facilitate multi-lingual research.

📄 PDF Abstract BibTeX arXiv:1703.00948

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesEntity EmbeddingsInformation RetrievalRetrieval

Similar Papers 제목 키워드 기반

WikiCoref: An English Coreference-annotated Corpus of Wikipedia Articles

2016-05-01 · LREC 2016 5 · Abbas Ghaddar, Phillippe Langlais

This paper presents WikiCoref, an English corpus annotated for anaphoric relations, where all documents are from the English version of Wikipedia. Our annotation scheme follows the one of OntoNotes with a few disparities…

Articlescoreference-resolutionCoreference Resolution

DBpedia Abstracts: A Large-Scale, Open, Multilingual NLP Training Corpus

2016-05-01 · LREC 2016 5 · Martin Br{\"u}mmer, Milan Dojchinovski, Sebastian Hellmann

The ever increasing importance of machine learning in Natural Language Processing is accompanied by an equally increasing need in large-scale training and evaluation corpora. Due to its size, its openness and relative qu…

Entity LinkingMultilingual NLP

Recognizing Descriptive Wikipedia Categories for Historical Figures

2017-04-24 · Yanqing Chen, Steven Skiena

Wikipedia is a useful knowledge source that benefits many applications in language processing and knowledge representation. An important feature of Wikipedia is that of categories. Wikipedia pages are assigned different …

DescriptiveInformation RetrievalRetrievalTAG

DDisCo: A Discourse Coherence Dataset for Danish

2022-06-01 · LREC 2022 6 · Linea Flansmose Mikkelsen, Oliver Kinch, Anders Jess Pedersen, Ophélie Lacroix

To date, there has been no resource for studying discourse coherence on real-world Danish texts. Discourse coherence has mostly been approached with the assumption that incoherent texts can be represented by coherent tex…

pioNER: Datasets and Baselines for Armenian Named Entity Recognition

2018-10-19 · Tsolak Ghukasyan, Garnik Davtyan, Karen Avetisyan, Ivan Andrianov

In this work, we tackle the problem of Armenian named entity recognition, providing silver- and gold-standard datasets as well as establishing baseline results on popular models. We present a 163000-token named entity co…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Word Embeddings