paper-with-me

홈 › Papers

Stochastic Tokenization with a Language Model for Neural Text Classification

2019-07-01 · ACL 2019 7 · Tatsuya Hiraoka, Hiroyuki Shindo, Yuji Matsumoto

For unsegmented languages such as Japanese and Chinese, tokenization of a sentence has a significant impact on the performance of text classification. Sentences are usually segmented with words or subwords by a morphological analyzer or byte pair encoding and then encoded with word (or subword) representations for neural networks. However, segmentation is potentially ambiguous, and it is unclear whether the segmented tokens achieve the best performance for the target task. In this paper, we propose a method to simultaneously learn tokenization and text classification to address these problems. Our model incorporates a language model for unsupervised tokenization into a text classifier and then trains both models simultaneously. To make the model robust against infrequent tokens, we sampled segmentation for each sentence stochastically during training, which resulted in improved performance of text classification. We conducted experiments on sentiment analysis as a text classification task and show that our method achieves better performance than previous methods.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGeneral ClassificationLanguage ModelingLanguage ModellingSegmentationSentenceSentiment Analysistext-classificationText Classification

Similar Papers 제목 키워드 기반

Evaluating Various Tokenizers for Arabic Text Classification

2021-06-14 · Zaid Alyafeai, Maged S. Al-shaibani, Mustafa Ghaleb, Irfan Ahmad

The first step in any NLP pipeline is to split the text into individual tokens. The most obvious and straightforward approach is to use words as tokens. However, given a large text corpus, representing all the words is n…

ClassificationNews ClassificationSentiment Analysistext-classification+1

Distributional Properties of Subword Regularization

2024-08-21 · Marco Cognetta, Vilém Zouhar, Naoaki Okazaki

Subword regularization, used widely in NLP, improves model performance by reducing the dependency on exact tokenizations, augmenting the training corpus, and exposing the model to more unique contexts during training. BP…

Machine Translation

Joint Optimization of Tokenization and Downstream Model

2021-05-26 · Findings (ACL) 2021 8 · Tatsuya Hiraoka, Sho Takase, Kei Uchiumi, Atsushi Keyaki 외

Since traditional tokenizers are isolated from a downstream task and model, they cannot output an appropriate tokenization depending on the task and model, although recent studies imply that the appropriate tokenization …

Machine Translationmodeltext-classificationText Classification+1

Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge

2024-04-20 · Khuyagbaatar Batsuren, Ekaterina Vylomova, Verna Dankers, Tsetsuukhei Delgerbaatar 외

The popular subword tokenizers of current language models, such as Byte-Pair Encoding (BPE), are known not to respect morpheme boundaries, which affects the downstream performance of the models. While many improved token…

text-classificationText Classification

Downstream Task-Oriented Neural Tokenizer Optimization with Vocabulary Restriction as Post Processing

2023-04-21 · Tatsuya Hiraoka, Tomoya Iwakura

This paper proposes a method to optimize tokenization for the performance improvement of already trained downstream models. Our method generates tokenization results attaining lower loss values of a given downstream mode…

text-classificationText Classification