paper-with-me

Papers

Wiktextract: Wiktionary as Machine-Readable Structured Data

2022-06-01 · LREC 2022 6 · Tatu Ylonen

We present a machine-readable structured data version of Wiktionary. Unlike previous Wiktionary extractions, the new extractor, Wiktextract, fully interprets and expands templates and Lua modules in Wiktionary. This enables it to perform a more complete, robust, and maintainable extraction. The extracted data is multilingual and includes lemmas, inflected forms, translations, etymology, usage examples, pronunciations (including URLs of sound files), lexical and semantic relations, and various morphological, syntactic, semantic, topical, and dialectal annotations. We extract all data from the English Wiktionary. Comparing against previous extractions from language-specific dictionaries, we find that its coverage for non-English languages often matches or exceeds the coverage in the language-specific editions, with the added benefit that all glosses are in English. The data is freely available and regularly updated, enabling anyone to add more data and correct errors by editing Wiktionary. The extracted data is in JSON format and designed to be easy to use by researchers, downstream resources, and application developers.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The comparison of Wiktionary thesauri transformed into the machine-readable format

2010-06-25 · A. A. Krizhanovsky

Wiktionary is a unique, peculiar, valuable and original resource for natural language processing (NLP). The paper describes an open-source Wiktionary parser: its architecture and requirements followed by a description of…

Multilingual ontology matching based on Wiktionary data accessible via SPARQL endpoint

2011-09-04 · Feiyu Lin, Andrew Krizhanovsky

Interoperability is a feature required by the Semantic Web. It is provided by the ontology matching methods and algorithms. But now ontologies are presented not only in English, but in other languages as well. It is impo…

Ontology MatchingTranslation

ENGLAWI: From Human- to Machine-Readable Wiktionary

2020-05-01 · LREC 2020 5 · Franck Sajous, Basilio Calderone, Nabil Hathout

This paper introduces ENGLAWI, a large, versatile, XML-encoded machine-readable dictionary extracted from Wiktionary. ENGLAWI contains 752,769 articles encoding the full body of information included in Wiktionary: simple…

ArticlesWord Embeddings

Analysis of the quotation corpus of the Russian Wiktionary

2020-01-20 · A. Smirnov, T. Levashova, A. Karpov, I. Kipyatkova 외

The quantitative evaluation of quotations in the Russian Wiktionary was performed using the developed Wiktionary parser. It was found that the number of quotations in the dictionary is growing fast (51.5 thousands in 201…

Wiktionnaire's Wikicode GLAWIfied: a Workable French Machine-Readable Dictionary

2016-05-01 · LREC 2016 5 · Nabil Hathout, Franck Sajous

GLAWI is a free, large-scale and versatile Machine-Readable Dictionary (MRD) that has been extracted from the French language edition of Wiktionary, called Wiktionnaire. In (Sajous and Hathout, 2015), we introduced GLAWI…

Diversity