The comparison of Wiktionary thesauri transformed into the machine-readable format
Wiktionary is a unique, peculiar, valuable and original resource for natural language processing (NLP). The paper describes an open-source Wiktionary parser: its architecture and requirements followed by a description of Wiktionary features to be taken into account, some open problems of Wiktionary and the parser. The current implementation of the parser extracts the definitions, semantic relations, and translations from English and Russian Wiktionaries. The paper's goal is to interest researchers (1) in using the constructed machine-readable dictionary for different NLP tasks, (2) in extending the software to parse 170 still unused Wiktionaries. The comparison of a number and types of semantic relations, a number of definitions, and a number of translations in the English Wiktionary and the Russian Wiktionary has been carried out. It was found that the number of semantic relations in the English Wiktionary is larger by 1.57 times than in Russian (157 and 100 thousands). But the Russian Wiktionary has more "rich" entries (with a big number of semantic relations), e.g. the number of entries with three or more semantic relations is larger by 1.63 times than in the English Wiktionary. Upon comparison, it was found out the methodological shortcomings of the Wiktionary.
Code (1)
Similar Papers 제목 키워드 기반
Multilingual ontology matching based on Wiktionary data accessible via SPARQL endpoint
Interoperability is a feature required by the Semantic Web. It is provided by the ontology matching methods and algorithms. But now ontologies are presented not only in English, but in other languages as well. It is impo…
Ontology MatchingTranslationExtrinsic Evaluation of French Dependency Parsers on a Specialized Corpus: Comparison of Distributional Thesauri
We present a study in which we compare 11 different French dependency parsers on a specialized corpus (consisting of research articles on NLP from the proceedings of the TALN conference). Due to the lack of a suitable go…
ArticlesCross-lingual RDF Thesauri Interlinking
Various lexical resources are being published in RDF. To enhance the usability of these resources, identical resources in different data sets should be linked. If lexical resources are described in different natural lang…
Machine TranslationTranslationWiktextract: Wiktionary as Machine-Readable Structured Data
We present a machine-readable structured data version of Wiktionary. Unlike previous Wiktionary extractions, the new extractor, Wiktextract, fully interprets and expands templates and Lua modules in Wiktionary. This enab…
Chinese Characters Mapping Table of Japanese, Traditional Chinese and Simplified Chinese
Chinese characters are used both in Japanese and Chinese, which are called Kanji and Hanzi respectively. Chinese characters contain significant semantic information, a mapping table between Kanji and Hanzi can be very us…
Cross-Lingual Information RetrievalInformation RetrievalMachine TranslationRelation+2