Comparing two analyzers of Japanese corpora for helping linguists: MeCab and Sagace (Comparaison de deux outils d'analyse de corpus japonais pour l'aide au linguiste, Sagace et Mecab) [in French]
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Shrinking Japanese Morphological Analyzers With Neural Networks and Semi-supervised Learning
For languages without natural word boundaries, like Japanese and Chinese, word segmentation is a prerequisite for downstream analysis. For Japanese, segmentation is often done jointly with part of speech tagging, and thi…
Chinese Word SegmentationMorphological AnalysisPart-Of-Speech TaggingSegmentationBenchmarking Large Language Models for Grapheme-to-Phoneme Conversion: A Japanese Case Study
Grapheme-to-phoneme (G2P) conversion is essential for controllable and robust text-to-speech, and large language models (LLMs), with broad linguistic knowledge, offer a promising approach. We benchmarked over 30 LLMs on …
JASS: Japanese-specific Sequence to Sequence Pre-training for Neural Machine Translation
Neural machine translation (NMT) needs large parallel corpora for state-of-the-art translation quality. Low-resource NMT is typically addressed by transfer learning which leverages large monolingual or parallel corpora f…
Low Resource NMTMachine TranslationNMTTransfer Learning+1Pre-training via Leveraging Assisting Languages for Neural Machine Translation
Sequence-to-sequence (S2S) pre-training using large monolingual data is known to improve performance for various S2S NLP tasks. However, large monolingual corpora might not always be available for the languages of intere…
Machine TranslationNMTTranslationBuilding a Japanese Typo Dataset from Wikipedia's Revision History
User generated texts contain many typos for which correction is necessary for NLP systems to work. Although a large number of typo{--}correction pairs are needed to develop a data-driven typo correction system, no such d…