Stochastic Tokenization with a Language Model for Neural Text Classification
For unsegmented languages such as Japanese and Chinese, tokenization of a sentence has a significant impact on the performance of text classification. Sentences are usually segmented with words or subwords by a morphological analyzer or byte pair encoding and then encoded with word (or subword) representations for neural networks. However, segmentation is potentially ambiguous, and it is unclear whether the segmented tokens achieve the best performance for the target task. In this paper, we propose a method to simultaneously learn tokenization and text classification to address these problems. Our model incorporates a language model for unsupervised tokenization into a text classifier and then trains both models simultaneously. To make the model robust against infrequent tokens, we sampled segmentation for each sentence stochastically during training, which resulted in improved performance of text classification. We conducted experiments on sentiment analysis as a text classification task and show that our method achieves better performance than previous methods.
Code (0)
등록된 구현이 없습니다.
Tasks
ClassificationGeneral ClassificationLanguage ModelingLanguage ModellingSegmentationSentenceSentiment Analysistext-classificationText ClassificationSimilar Papers 제목 키워드 기반
Evaluating Various Tokenizers for Arabic Text Classification
The first step in any NLP pipeline is to split the text into individual tokens. The most obvious and straightforward approach is to use words as tokens. However, given a large text corpus, representing all the words is n…
ClassificationNews ClassificationSentiment Analysistext-classification+1Distributional Properties of Subword Regularization
Subword regularization, used widely in NLP, improves model performance by reducing the dependency on exact tokenizations, augmenting the training corpus, and exposing the model to more unique contexts during training. BP…
Machine TranslationJoint Optimization of Tokenization and Downstream Model
Since traditional tokenizers are isolated from a downstream task and model, they cannot output an appropriate tokenization depending on the task and model, although recent studies imply that the appropriate tokenization …
Machine Translationmodeltext-classificationText Classification+1Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
The popular subword tokenizers of current language models, such as Byte-Pair Encoding (BPE), are known not to respect morpheme boundaries, which affects the downstream performance of the models. While many improved token…
text-classificationText ClassificationDownstream Task-Oriented Neural Tokenizer Optimization with Vocabulary Restriction as Post Processing
This paper proposes a method to optimize tokenization for the performance improvement of already trained downstream models. Our method generates tokenization results attaining lower loss values of a given downstream mode…
text-classificationText Classification