Breaking Character: Are Subwords Good Enough for MRLs After All?
Large pretrained language models (PLMs) typically tokenize the input string into contiguous subwords before any pretraining or inference. However, previous studies have claimed that this form of subword tokenization is inadequate for processing morphologically-rich languages (MRLs). We revisit this hypothesis by pretraining a BERT-style masked language model over character sequences instead of word-pieces. We compare the resulting model, dubbed TavBERT, against contemporary PLMs based on subwords for three highly complex and ambiguous MRLs (Hebrew, Turkish, and Arabic), testing them on both morphological and semantic tasks. Our results show, for all tested languages, that while TavBERT obtains mild improvements on surface-level tasks à la POS tagging and full morphological disambiguation, subword-based PLMs achieve significantly higher performance on semantic tasks, such as named entity recognition and extractive question answering. These results showcase and (re)confirm the potential of subword tokenization as a reasonable modeling assumption for many languages, including MRLs.
Code (0)
등록된 구현이 없습니다.
Tasks
AllExtractive Question-AnsweringLanguage ModelingLanguage ModellingMorphological Disambiguationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)POSPOS TaggingQuestion AnsweringSimilar Papers 제목 키워드 기반
Breaking Character: Are Subwords Good Enough for MRLs After All?
Large pretrained language models (PLMs) typically tokenize the input string into contiguous subwords before any pretraining or inference. However, previous studies have claimed that this form of subword tokenization is i…
AllExtractive Question-AnsweringLanguage ModelingLanguage Modelling+7Optimal Turkish Subword Strategies at Scale: Systematic Evaluation of Data, Vocabulary, Morphology Interplay
Tokenization is a pivotal design choice for neural language modeling in morphologically rich languages (MRLs) such as Turkish, where productive agglutination challenges both vocabulary efficiency and morphological fideli…
Dependency ParsingSentiment AnalysisTailoring Neural Architectures for Translating from Morphologically Rich Languages
A morphologically complex word (MCW) is a hierarchical constituent with meaning-preserving subunits, so word-based models which rely on surface forms might not be powerful enough to translate such structures. When transl…
DecoderMachine TranslationNMTSentence+1Improving Character-based Decoding Using Target-Side Morphological Information for Neural Machine Translation
Recently, neural machine translation (NMT) has emerged as a powerful alternative to conventional statistical approaches. However, its performance drops considerably in the presence of morphologically rich languages (MRLs…
DecoderMachine TranslationNMTTranslationOn the Impact of the Cutoff Time on the Performance of Algorithm Configurators
Algorithm configurators are automated methods to optimise the parameters of an algorithm for a class of problems. We evaluate the performance of a simple random local search configurator (ParamRLS) for tuning the neighbo…