Corpus-based Check-up for Thesaurus
In this paper we discuss the usefulness of applying a checking procedure to existing thesauri. The procedure is based on the analysis of discrepancies of corpus-based and thesaurus-based word similarities. We applied the procedure to more than 30 thousand words of the Russian wordnet and found some serious errors in word sense description, including inaccurate relationships and missing senses of ambiguous words.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Corpus+WordNet thesaurus generation for ontology enriching
This paper presents a model to enrich an ontology with a thesaurus based on a domain corpus and WordNet. The model is applied to the data privacy domain and the initial domain resources comprise a data privacy ontology, …
Semantic Textual SimilarityA spell-checker and thesaurus for Bambara (Bamanankan) (Un v\'erificateur orthographique pour la langue bambara) [in French]
Towards Automatic Thesaurus Construction and Enrichment.
Thesaurus construction with minimum human efforts often relies on automatic methods to discover terms and their relations. Hence, the quality of a thesaurus heavily depends on the chosen methodologies for: (i) building i…
Semantic SimilaritySemantic Textual SimilarityDistributed Distributional Similarities of Google Books Over the Centuries
This paper introduces a distributional thesaurus and sense clusters computed on the complete Google Syntactic N-grams, which is extracted from Google Books, a very large corpus of digitized books published between 1520 a…
Graph ClusteringThesaurus Verification Based on Distributional Similarities
In this paper we consider an approach to verification of large lexical-semantic resources as WordNet. The method of verification procedure is based on the analysis of discrepancies of corpus-based and thesaurus-based wor…