IndoCollex: A Testbed for Morphological Transformation of Indonesian Word Colloquialism
Code (1)
Similar Papers 제목 키워드 기반
IDENTIC Corpus: Morphologically Enriched Indonesian-English Parallel Corpus
This paper describes the creation process of an Indonesian-English parallel corpus (IDENTIC). The corpus contains 45,000 sentences collected from different sources in different genres. Several manual text preprocessing t…
Spelling CorrectionFine-tuning Pretrained Multilingual BERT Model for Indonesian Aspect-based Sentiment Analysis
Although previous research on Aspect-based Sentiment Analysis (ABSA) for Indonesian reviews in hotel domain has been conducted using CNN and XGBoost, its model did not generalize well in test data and high number of OOV …
Aspect-Based Sentiment AnalysisAspect-Based Sentiment Analysis (ABSA)Sentiment AnalysisA Neural Network Approach to Create Minangkabau-Indonesia Bilingual Dictionary
Indonesia has many varieties of ethnic languages, and most come from the same language family, namely Austronesian languages. Coming from that same language family, the words in Indonesian ethnic languages are very simil…
DecoderKaWAT: A Word Analogy Task Dataset for Indonesian
We introduced KaWAT (Kata Word Analogy Task), a new word analogy task dataset for Indonesian. We evaluated on it several existing pretrained Indonesian word embeddings and embeddings trained on Indonesian online news cor…
Word EmbeddingsExploiting Morphological Regularities in Distributional Word Representations
We present an unsupervised, language agnostic approach for exploiting morphological regularities present in high dimensional vector spaces. We propose a novel method for generating embeddings of words from their morpholo…
ChunkingDocument ClassificationQuestion AnsweringWord Embeddings