paper-with-me

홈 › Papers

An Experimental Evaluation of Japanese Tokenizers for Sentiment-Based Text Classification

2024-12-23 · Andre Rusli, Makoto Shishido

This study investigates the performance of three popular tokenization tools: MeCab, Sudachi, and SentencePiece, when applied as a preprocessing step for sentiment-based text classification of Japanese texts. Using Term Frequency-Inverse Document Frequency (TF-IDF) vectorization, we evaluate two traditional machine learning classifiers: Multinomial Naive Bayes and Logistic Regression. The results reveal that Sudachi produces tokens closely aligned with dictionary definitions, while MeCab and SentencePiece demonstrate faster processing speeds. The combination of SentencePiece, TF-IDF, and Logistic Regression outperforms the other alternatives in terms of classification performance.

📄 PDF Abstract BibTeX arXiv:2412.17361

Code (1)

arusl/anlp_nlp2021_d3-1 공식 구현

Tasks

regressiontext-classificationText Classification

Methods 이 논문이 사용한 방법론

BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
SentencePiece 설명 없음
Logistic Regression Logistic Regression, despite its name, is a linear model for classification rather than regression. Logistic regression is also known in the literature as logit regression,…

Similar Papers 제목 키워드 기반

How do different tokenizers perform on downstream tasks in scriptio continua languages?: A case study in Japanese

2023-06-16 · Takuro Fujii, Koki Shibata, Atsuki Yamaguchi, Terufumi Morishita 외

This paper investigates the effect of tokenizers on the downstream performance of pretrained language models (PLMs) in scriptio continua languages where no explicit spaces exist between words, using Japanese as a case st…

Japanese Sentiment Classification using a Tree-Structured Long Short-Term Memory with Attention

2017-04-04 · PACLIC 2018 12 · Ryosuke Miyazaki, Mamoru Komachi

Previous approaches to training syntax-based sentiment classification models required phrase-level annotated corpora, which are not readily available in many languages other than English. Thus, we propose the use of tree…

ClassificationGeneral ClassificationSentiment AnalysisSentiment Classification

JFinTEB: Japanese Financial Text Embedding Benchmark

2026-04-17 · Masahiro Suzuki, Hiroki Sakaji arxiv

We introduce JFinTEB, the first comprehensive benchmark specifically designed for evaluating Japanese financial text embeddings. Existing embedding benchmarks provide limited coverage of language-specific and domain-spec…

Sentiment AnalysisText Generation

fugashi, a Tool for Tokenizing Japanese in Python

2020-10-14 · EMNLP (NLPOSS) 2020 11 · Paul McCann

Recent years have seen an increase in the number of large-scale multilingual NLP projects. However, even in such projects, languages with special processing requirements are often excluded. One such language is Japanese.…

Multilingual NLP

SiSP: Japanese Situation-dependent Sentiment Polarity Dictionary

2021-11-16 · ACL ARR November 2021 11 · Anonymous

In order to deal with the variety of meanings and contexts of words, we created a Japanese Situation-dependent Sentiment Polarity Dictionary (SiSP) of sentiment values labeled for 20 different situations. This dictionary…