paper-with-me

Papers

Identifying and Exploiting Definitions in Wordnet Bahasa

2016-01-01 · GWC 2016 1 · David Moeljadi, Francis Bond

This paper describes our attempts to add Indonesian definitions to synsets in the Wordnet Bahasa (Nurril Hirfana Mohamed Noor et al., 2011; Bond et al., 2014), to extract semantic relations between lemmas and definitions for nouns and verbs, such as synonym, hyponym, hypernym and instance hypernym, and to generally improve Wordnet. The original, somewhat noisy, definitions for Indonesian came from the Asian Wordnet project (Riza et al., 2010). The basic method of extracting the relations is based on Bond et al. (2004). Before the relations can be extracted, the definitions were cleaned up and tokenized. We found that the definitions cannot be completely cleaned up because of many misspellings and bad translations. However, we could identify four semantic relations in 57.10% of noun and verb definitions. For the remaining 42.90%, we propose to add 149 new Indonesian lemmas and make some improvements to Wordnet Bahasa and Wordnet in general.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Categorization of Semantic Roles for Dictionary Definitions

2018-06-20 · WS 2016 12 · Vivian S. Silva, Siegfried Handschuh, André Freitas

Understanding the semantic relationships between terms is a fundamental task in natural language processing applications. While structured resources that can express those relationships in a formal way, such as ontologie…

Categorization of Semantic Roles for Dictionary Definitions

2016-12-01 · WS 2016 12 · Vivian Silva, H, Siegfried schuh, Andr{\'e} Freitas

Understanding the semantic relationships between terms is a fundamental task in natural language processing applications. While structured resources that can express those relationships in a formal way, such as ontologie…

Question Answering

Building a Knowledge Graph from Natural Language Definitions for Interpretable Text Entailment Recognition

2018-06-20 · LREC 2018 5 · Vivian S. Silva, André Freitas, Siegfried Handschuh

Natural language definitions of terms can serve as a rich source of knowledge, but structuring them into a comprehensible semantic model is essential to enable them to be used in semantic interpretation tasks. We propose…

World Knowledge

BRCC and SentiBahasaRojak: The First Bahasa Rojak Corpus for Pretraining and Sentiment Analysis Dataset

2022-10-01 · COLING 2022 10 · Nanda Putri Romadhona, Sin-En Lu, Bo-Han Lu, Richard Tzong-Han Tsai

Code-mixing refers to the mixed use of multiple languages. It is prevalent in multilingual societies and is also one of the most challenging natural language processing tasks. In this paper, we study Bahasa Rojak, a dial…

Data AugmentationSentiment AnalysisTAG

CILI: the Collaborative Interlingual Index

2016-01-01 · GWC 2016 1 · Francis Bond, Piek Vossen, John McCrae, Christiane Fellbaum

This paper introduces the motivation for and design of the Collaborative InterLingual Index (CILI). It is designed to make possible coordination between multiple loosely coupled wordnet projects. The structure of the CIL…