Automatically Extracting Variant-Normalization Pairs for Japanese Text Normalization
Social media texts, such as tweets from Twitter, contain many types of non-standard tokens, and the number of normalization approaches for handling such noisy text has been increasing. We present a method for automatically extracting pairs of a variant word and its normal form from unsegmented text on the basis of a pair-wise similarity approach. We incorporated the acquired variant-normalization pairs into Japanese morphological analysis. The experimental results show that our method can extract widely covered variants from large Twitter data and improve the recall of normalization without degrading the overall accuracy of Japanese morphological analysis.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationMorphological AnalysisText NormalizationSimilar Papers 제목 키워드 기반
Japanese Text Normalization with Encoder-Decoder Model
Text normalization is the task of transforming lexical variants to their canonical forms. We model the problem of text normalization as a character-level sequence to sequence learning problem and present a neural encoder…
Data AugmentationDecoderMachine Translationmodel+3Are Girls Neko or Shōjo? Cross-Lingual Alignment of Non-Isomorphic Embeddings with Iterative Normalization
Cross-lingual word embeddings (CLWE) underlie many multilingual natural language processing systems, often through orthogonal transformations of pre-trained monolingual embeddings. However, orthogonal mapping only works …
Cross-Lingual Word EmbeddingsTranslationWord EmbeddingsWord TranslationAre Girls Neko or Sh\=ojo? Cross-Lingual Alignment of Non-Isomorphic Embeddings with Iterative Normalization
Cross-lingual word embeddings (CLWE) underlie many multilingual natural language processing systems, often through orthogonal transformations of pre-trained monolingual embeddings. However, orthogonal mapping only works …
Cross-Lingual Word EmbeddingsTranslationWord EmbeddingsWord TranslationFindings of the WMT 2020 Shared Task on Machine Translation Robustness
We report the findings of the second edition of the shared task on improving robustness in Machine Translation (MT). The task aims to test current machine translation systems in their ability to handle challenges facing …
DiversityMachine TranslationTranslationUser-Generated Text Corpus for Evaluating Japanese Morphological Analysis and Lexical Normalization
Morphological analysis (MA) and lexical normalization (LN) are both important tasks for Japanese user-generated text (UGT). To evaluate and compare different MA/LN systems, we have constructed a publicly available Japane…
Lexical NormalizationMorphological Analysis