Unsupervised morph segmentation and statistical language models for vocabulary expansion
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech Recognition (ASR)Language ModelingLanguage ModellingMachine TranslationMORPHOptical Character Recognition (OCR)Speech RecognitionSimilar Papers 제목 키워드 기반
Effects of sub-word segmentation on performance of transformer language models
Language modeling is a fundamental task in natural language processing, which has been thoroughly explored with various architectures and hyperparameters. However, few studies focus on the effect of sub-word segmentation…
Language ModelingLanguage ModellingSegmentationUnsupervised Morphological Tree Tokenizer
As a cornerstone in language modeling, tokenization involves segmenting text inputs into pre-defined atomic units. Conventional statistical tokenizers often disrupt constituent boundaries within words, thereby corrupting…
Language ModelingLanguage ModellingSubword Segmental Language Modelling for Nguni Languages
Subwords have become the standard units of text in NLP, enabling efficient open-vocabulary models. With algorithms like byte-pair encoding (BPE), subword segmentation is viewed as a preprocessing step applied to the corp…
Language ModelingLanguage ModellingSegmentationLinguistically Motivated Vocabulary Reduction for Neural Machine Translation from Turkish to English
The necessity of using a fixed-size word vocabulary in order to control the model complexity in state-of-the-art neural machine translation (NMT) systems is an important bottleneck on performance, especially for morpholo…
Machine TranslationMorphological AnalysisNMTTranslationParsimonious Morpheme Segmentation with an Application to Enriching Word Embeddings
Traditionally, many text-mining tasks treat individual word-tokens as the finest meaningful semantic granularity. However, in many languages and specialized corpora, words are composed by concatenating semantically meani…
Language ModelingLanguage ModellingSegmentationWord Embeddings