paper-with-me

Papers

Substring Frequency Features for Segmentation of Japanese Katakana Words with Unlabeled Corpora

2017-11-01 · IJCNLP 2017 11 · Yoshinari Fujinuma, Alvin Grissom II

Word segmentation is crucial in natural language processing tasks for unsegmented languages. In Japanese, many out-of-vocabulary words appear in the phonetic syllabary katakana, making segmentation more difficult due to the lack of clues found in mixed script settings. In this paper, we propose a straightforward approach based on a variant of tf-idf and apply it to the problem of word segmentation in Japanese. Even though our method uses only an unlabeled corpus, experimental results show that it achieves performance comparable to existing methods that use manually labeled corpora. Furthermore, it improves performance of simple word segmentation models trained on a manually labeled corpus.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalMachine TranslationSegmentation

Similar Papers 제목 키워드 기반

A Comparison of Entity Matching Methods between English and Japanese Katakana

2018-10-01 · WS 2018 10 · Michiharu Yamashita, Hideki Awashima, Hidekazu Oiwa

Japanese Katakana is one component of the Japanese writing system and is used to express English terms, loanwords, and onomatopoeia in Japanese characters based on the phonemes. The main purpose of this research is to fi…

Transliteration

Long Short-Term Memory for Japanese Word Segmentation

2017-09-23 · PACLIC 2018 12 · Yoshiaki Kitagawa, Mamoru Komachi

This study presents a Long Short-Term Memory (LSTM) neural network approach to Japanese word segmentation (JWS). Previous studies on Chinese word segmentation (CWS) succeeded in using recurrent neural networks such as LS…

Chinese Word SegmentationJapanese Word SegmentationSegmentation

Dictionary Look-up with Katakana Variant Recognition

2012-05-01 · LREC 2012 5 · Satoshi Sato

The Japanese language has rich variety and quantity of word variant. Since 1980s, it has been recognized that this richness becomes an obstacle against natural language processing. A complete solution, however, has not b…

Morphological AnalysisRetrievalTransliteration

Compact and Robust Models for Japanese-English Character-level Machine Translation

2019-11-01 · WS 2019 11 · Jinan Dai, Kazunori Yamaguchi

Character-level translation has been proved to be able to achieve preferable translation quality without explicit segmentation, but training a character-level model needs a lot of hardware resources. In this paper, we in…

Machine TranslationTranslation

Memory-Efficient Katakana Compound Segmentation using Conditional Random Fields

2012-12-01 · COLING 2012 12 · Krauchanka Siarhei, Artsimenya Artsiom
Morphological Analysis