ViLexNorm: A Lexical Normalization Corpus for Vietnamese Social Media Text
Lexical normalization, a fundamental task in Natural Language Processing (NLP), involves the transformation of words into their canonical forms. This process has been proven to benefit various downstream NLP tasks greatly. In this work, we introduce Vietnamese Lexical Normalization (ViLexNorm), the first-ever corpus developed for the Vietnamese lexical normalization task. The corpus comprises over 10,000 pairs of sentences meticulously annotated by human annotators, sourced from public comments on Vietnam's most popular social media platforms. Various methods were used to evaluate our corpus, and the best-performing system achieved a result of 57.74% using the Error Reduction Rate (ERR) metric (van der Goot, 2019a) with the Leave-As-Is (LAI) baseline. For extrinsic evaluation, employing the model trained on ViLexNorm demonstrates the positive impact of the Vietnamese lexical normalization task on other NLP tasks. Our corpus is publicly available exclusively for research purposes.
Code (1)
Tasks
Lexical NormalizationVietnamese Lexical NormalizationVietnamese Social Media Text ProcessingSimilar Papers 제목 키워드 기반
ViSoLex: An Open-Source Repository for Vietnamese Social Media Lexical Normalization
ViSoLex is an open-source system designed to address the unique challenges of lexical normalization for Vietnamese social media text. The platform provides two core services: Non-Standard Word (NSW) Lookup and Lexical No…
Lexical NormalizationWeakly-supervised LearningA Weakly Supervised Data Labeling Framework for Machine Lexical Normalization in Vietnamese Social Media
This study introduces an innovative automatic labeling framework to address the challenges of lexical normalization in social media texts for low-resource languages like Vietnamese. Social media data is rich and diverse,…
Lexical NormalizationViGoEmotions: A Benchmark Dataset For Fine-grained Emotion Detection on Vietnamese Texts
Emotion classification plays a significant role in emotion prediction and harmful content detection. Recent advancements in NLP, particularly through large language models (LLMs), have greatly improved outcomes in this f…
Emotion ClassificationLexical NormalizationLexical Normalization of User-Generated Medical Text
In the medical domain, user-generated social media text is increasingly used as a valuable complementary knowledge source to scientific medical literature. The extraction of this knowledge is complicated by colloquial la…
Lexical NormalizationMistake DetectionSpelling CorrectionAutomatic Textual Normalization for Hate Speech Detection
Social media data is a valuable resource for research, yet it contains a wide range of non-standard words (NSW). These irregularities hinder the effective operation of NLP tools. Current state-of-the-art methods for the …
Hate Speech DetectionLexical NormalizationVietnamese Hate Speech Detection