paper-with-me

홈 › Papers

Comparing two analyzers of Japanese corpora for helping linguists: MeCab and Sagace (Comparaison de deux outils d'analyse de corpus japonais pour l'aide au linguiste, Sagace et Mecab) [in French]

2014-07-01 · JEPTALNRECITAL 2014 7 · Raoul Blin
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Shrinking Japanese Morphological Analyzers With Neural Networks and Semi-supervised Learning

2019-06-01 · NAACL 2019 6 · Arseny Tolmachev, Daisuke Kawahara, Sadao Kurohashi

For languages without natural word boundaries, like Japanese and Chinese, word segmentation is a prerequisite for downstream analysis. For Japanese, segmentation is often done jointly with part of speech tagging, and thi…

Chinese Word SegmentationMorphological AnalysisPart-Of-Speech TaggingSegmentation

Benchmarking Large Language Models for Grapheme-to-Phoneme Conversion: A Japanese Case Study

2026-06-20 · Tomoki Koriyama arxiv

Grapheme-to-phoneme (G2P) conversion is essential for controllable and robust text-to-speech, and large language models (LLMs), with broad linguistic knowledge, offer a promising approach. We benchmarked over 30 LLMs on …

JASS: Japanese-specific Sequence to Sequence Pre-training for Neural Machine Translation

2020-05-07 · LREC 2020 5 · Zhuoyuan Mao, Fabien Cromieres, Raj Dabre, Haiyue Song 외

Neural machine translation (NMT) needs large parallel corpora for state-of-the-art translation quality. Low-resource NMT is typically addressed by transfer learning which leverages large monolingual or parallel corpora f…

Low Resource NMTMachine TranslationNMTTransfer Learning+1

Pre-training via Leveraging Assisting Languages for Neural Machine Translation

2020-07-01 · ACL 2020 6 · Haiyue Song, Raj Dabre, Zhuoyuan Mao, Fei Cheng 외

Sequence-to-sequence (S2S) pre-training using large monolingual data is known to improve performance for various S2S NLP tasks. However, large monolingual corpora might not always be available for the languages of intere…

Machine TranslationNMTTranslation

Building a Japanese Typo Dataset from Wikipedia's Revision History

2020-07-01 · ACL 2020 6 · Yu Tanaka, Yugo Murawaki, Daisuke Kawahara, Sadao Kurohashi

User generated texts contain many typos for which correction is necessary for NLP systems to work. Although a large number of typo{--}correction pairs are needed to develop a data-driven typo correction system, no such d…