paper-with-me

홈 › Papers

Breaking Character: Are Subwords Good Enough for MRLs After All?

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Large pretrained language models (PLMs) typically tokenize the input string into contiguous subwords before any pretraining or inference. However, previous studies have claimed that this form of subword tokenization is inadequate for processing morphologically-rich languages (MRLs). We revisit this hypothesis by pretraining a BERT-style masked language model over character sequences instead of word-pieces. We compare the resulting model, dubbed TavBERT, against contemporary PLMs based on subwords for three highly complex and ambiguous MRLs (Hebrew, Turkish, and Arabic), testing them on both morphological and semantic tasks. Our results show, for all tested languages, that while TavBERT obtains mild improvements on surface-level tasks à la POS tagging and full morphological disambiguation, subword-based PLMs achieve significantly higher performance on semantic tasks, such as named entity recognition and extractive question answering. These results showcase and (re)confirm the potential of subword tokenization as a reasonable modeling assumption for many languages, including MRLs.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

AllExtractive Question-AnsweringLanguage ModelingLanguage ModellingMorphological Disambiguationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)POSPOS TaggingQuestion Answering

Similar Papers 제목 키워드 기반

Breaking Character: Are Subwords Good Enough for MRLs After All?

2022-04-10 · Omri Keren, Tal Avinari, Reut Tsarfaty, Omer Levy

Large pretrained language models (PLMs) typically tokenize the input string into contiguous subwords before any pretraining or inference. However, previous studies have claimed that this form of subword tokenization is i…

AllExtractive Question-AnsweringLanguage ModelingLanguage Modelling+7

Optimal Turkish Subword Strategies at Scale: Systematic Evaluation of Data, Vocabulary, Morphology Interplay

2026-02-06 · Duygu Altinok arxiv

Tokenization is a pivotal design choice for neural language modeling in morphologically rich languages (MRLs) such as Turkish, where productive agglutination challenges both vocabulary efficiency and morphological fideli…

Dependency ParsingSentiment Analysis

Tailoring Neural Architectures for Translating from Morphologically Rich Languages

2018-08-01 · COLING 2018 8 · Peyman Passban, Andy Way, Qun Liu

A morphologically complex word (MCW) is a hierarchical constituent with meaning-preserving subunits, so word-based models which rely on surface forms might not be powerful enough to translate such structures. When transl…

DecoderMachine TranslationNMTSentence+1

Improving Character-based Decoding Using Target-Side Morphological Information for Neural Machine Translation

2018-04-17 · NAACL 2018 6 · Peyman Passban, Qun Liu, Andy Way

Recently, neural machine translation (NMT) has emerged as a powerful alternative to conventional statistical approaches. However, its performance drops considerably in the presence of morphologically rich languages (MRLs…

DecoderMachine TranslationNMTTranslation

On the Impact of the Cutoff Time on the Performance of Algorithm Configurators

2019-04-12 · George T. Hall, Pietro S. Oliveto, Dirk Sudholt

Algorithm configurators are automated methods to optimise the parameters of an algorithm for a class of problems. We evaluate the performance of a simple random local search configurator (ParamRLS) for tuning the neighbo…