paper-with-me

홈 › Papers

A Study on the Appropriate size of the Mongolian general corpus

2023-07-12 · Sunsoo Choi, Ganbat Tsend

This study aims to determine the appropriate size of the Mongolian general corpus. This study used the Heaps function and Type Token Ratio to determine the appropriate size of the Mongolian general corpus. The sample corpus of 906,064 tokens comprised texts from 10 domains of newspaper politics, economy, society, culture, sports, world articles and laws, middle and high school literature textbooks, interview articles, and podcast transcripts. First, we estimated the Heaps function with this sample corpus. Next, we observed changes in the number of types and TTR values while increasing the number of tokens by one million using the estimated Heaps function. As a result of observation, we found that the TTR value hardly changed when the number of tokens exceeded from 39 to 42 million. Thus, we conclude that an appropriate size for a Mongolian general corpus is from 39 to 42 million tokens.

📄 PDF Abstract BibTeX arXiv:2307.06050

Code (0)

등록된 구현이 없습니다.

Tasks

Articles

Similar Papers 제목 키워드 기반

Mongolian Named Entity Recognition System with Rich Features

2016-12-01 · COLING 2016 12 · Weihua Wang, Feilong Bao, Guanglai Gao

In this paper, we first build a manually annotated named entity corpus of Mongolian. Then, we propose three morphological processing methods and study comprehensive features, including syllable features, lexical features…

Machine Translationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

Mongolian Questions Classification Based on Mulit-Head Attention

2020-10-01 · CCL 2020 10 · Guangyi Wang, Feilong Bao, Weihua Wang

Question classification is a crucial subtask in question answering system. Mongolian is a kind of few resource language. It lacks public labeled corpus. And the complex morphological structure of Mongolian vocabulary mak…

ClassificationQuestion Answering

A LSTM Approach with Sub-Word Embeddings for Mongolian Phrase Break Prediction

2018-08-01 · COLING 2018 8 · Rui Liu, Feilong Bao, Guanglai Gao, HUI ZHANG 외

In this paper, we first utilize the word embedding that focuses on sub-word units to the Mongolian Phrase Break (PB) prediction task by using Long-Short-Term-Memory (LSTM) model. Mongolian is an agglutinative language. E…

Dictionary LearningMachine TranslationQuestion AnsweringWord Embeddings

Interactive Mongolian Question Answer Matching Model Based on Attention Mechanism in the Law Domain

2022-10-01 · CCL 2022 10 · Peng Yutao, Wang Weihua, Bao Feilong

“Mongolian question answer matching task is challenging, since Mongolian is a kind of lowresource language and its complex morphological structures lead to data sparsity. In this work, we propose an Interactive Mongolian…

Question Answering

Deep learning model for Mongolian Citizens Feedback Analysis using Word Vector Embeddings

2023-02-23 · Zolzaya Dashdorj, Tsetsentsengel Munkhbayar, Stanislav Grigorev

A large amount of feedback was collected over the years. Many feedback analysis models have been developed focusing on the English language. Recognizing the concept of feedback is challenging and crucial in languages whi…

Deep LearningSentenceWord Embeddings