paper-with-me

Papers

Multi-Fusion Chinese WordNet (MCW) : Compound of Machine Learning and Manual Correction

2020-02-05 · Mingchen Li, Zili Zhou, Yanna Wang

Princeton WordNet (PWN) is a lexicon-semantic network based on cognitive linguistics, which promotes the development of natural language processing. Based on PWN, five Chinese wordnets have been developed to solve the problems of syntax and semantics. They include: Northeastern University Chinese WordNet (NEW), Sinica Bilingual Ontological WordNet (BOW), Southeast University Chinese WordNet (SEW), Taiwan University Chinese WordNet (CWN), Chinese Open WordNet (COW). By using them, we found that these word networks have low accuracy and coverage, and cannot completely portray the semantic network of PWN. So we decided to make a new Chinese wordnet called Multi-Fusion Chinese Wordnet (MCW) to make up those shortcomings. The key idea is to extend the SEW with the help of Oxford bilingual dictionary and Xinhua bilingual dictionary, and then correct it. More specifically, we used machine learning and manual adjustment in our corrections. Two standards were formulated to help our work. We conducted experiments on three tasks including relatedness calculation, word similarity and word sense disambiguation for the comparison of lemma's accuracy, at the same time, coverage also was compared. The results indicate that MCW can benefit from coverage and accuracy via our method. However, it still has room for improvement, especially with lemmas. In the future, we will continue to enhance the accuracy of MCW and expand the concepts in it.

📄 PDF Abstract BibTeX arXiv:2002.01761

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningWord Sense DisambiguationWord Similarity

Similar Papers 제목 키워드 기반

A Machine Learning Approach for the Identification of Bengali Noun-Noun Compound Multiword Expressions

2014-01-25 · Vivekananda Gayen, Kamal Sarkar

This paper presents a machine learning approach for identification of Bengali multiword expressions (MWE) which are bigram nominal compounds. Our proposed approach has two steps: (1) candidate extraction using chunk info…

BIG-bench Machine Learning

Sinitic Wordnet: Laying the Groundwork with Chinese Varieties Written in Traditional Characters

2018-01-01 · GWC 2018 1 · Chih-Yao Lee, Shu-Kai Hsieh

The present work seeks to make the logographic nature of Chinese script a relevant research ground in wordnet studies. While wordnets are not so much about words as about the concepts represented in words, synset formati…

Detection of Compound Nouns and Light Verb Constructions using IndoWordNet

2016-01-01 · GWC 2016 1 · Dhirendra Singh, Sudha Bhingardive, Pushpak Bhattacharyyaa

Detection of MultiWord Expressions (MWEs) is one of the fundamental problems in Natural Language Processing. In this paper, we focus on two categories of MWEs - Compound Nouns and Light Verb Constructions. These two cate…

Samāsa-Kartā: An Online Tool for Producing Compound Words using IndoWordNet

2016-01-01 · GWC 2016 1 · Hanumant Redkar, Nilesh Joshi, Sandhya Singh, Irawati Kulkarni 외

Samāsa or compounds are a regular feature of Indian Languages. They are also found in other languages like German, Italian, French, Russian, Spanish, etc. Compound word is constructed from two or more words to form a sin…

Morphological Analysis

Towards linking synonymous expressions of compound verbs to Japanese WordNet

2019-07-01 · GWC 2019 7 · Kyoko Kanzaki, Hitoshi Isahara

This paper describes our project on Japanese compound verbs. Japanese “Verb (adnominal form) + Verb” compounds, which are treated as single verbs, frequently appear in daily communication. They are not sufficiently regis…