Papers Lexical Normalization
“Lexical Normalization” 태그가 달린 논문 49편 · 필터 해제
ViGoEmotions: A Benchmark Dataset For Fine-grained Emotion Detection on Vietnamese Texts
Emotion classification plays a significant role in emotion prediction and harmful content detection. Recent advancements in NLP, particularly through large language models (LLMs), have greatly improved outcomes in this f…
Emotion ClassificationLexical NormalizationMultiLexNorm++: A Unified Benchmark and a Generative Model for Lexical Normalization for Asian Languages
Social media data has been of interest to Natural Language Processing (NLP) practitioners for over a decade, because of its richness in information, but also challenges for automatic processing. Since language use is mor…
Lexical NormalizationViSoLex: An Open-Source Repository for Vietnamese Social Media Lexical Normalization
ViSoLex is an open-source system designed to address the unique challenges of lexical normalization for Vietnamese social media text. The platform provides two core services: Non-Standard Word (NSW) Lookup and Lexical No…
Lexical NormalizationWeakly-supervised LearningA Weakly Supervised Data Labeling Framework for Machine Lexical Normalization in Vietnamese Social Media
This study introduces an innovative automatic labeling framework to address the challenges of lexical normalization in social media texts for low-resource languages like Vietnamese. Social media data is rich and diverse,…
Lexical NormalizationViLexNorm: A Lexical Normalization Corpus for Vietnamese Social Media Text
Lexical normalization, a fundamental task in Natural Language Processing (NLP), involves the transformation of words into their canonical forms. This process has been proven to benefit various downstream NLP tasks greatl…
Lexical NormalizationVietnamese Lexical NormalizationVietnamese Social Media Text ProcessingAutomatic Textual Normalization for Hate Speech Detection
Social media data is a valuable resource for research, yet it contains a wide range of non-standard words (NSW). These irregularities hinder the effective operation of NLP tools. Current state-of-the-art methods for the …
Hate Speech DetectionLexical NormalizationVietnamese Hate Speech DetectionIncreasing Robustness for Cross-domain Dialogue Act Classification on Social Media Data
Automatically detecting the intent of an utterance is important for various downstream natural language processing tasks. This task is also called Dialogue Act Classification (DAC) and was primarily researched on spoken …
Dialogue Act ClassificationLexical NormalizationA Character-level Ngram-based MT Approach for Lexical Normalization in Social Media
This paper presents an ngram-based MT approach that operates at character-level to generate possible canonical forms for lexical variants in social media text. It utilizes a joint n-gram model to learn edit sequences of …
Lexical NormalizationA Text Editing Approach to Joint Japanese Word Segmentation, POS Tagging, and Lexical Normalization
Lexical normalization, in addition to word segmentation and part-of-speech tagging, is a fundamental task for Japanese user-generated text processing. In this paper, we propose a text editing model to solve the three tas…
Japanese Word SegmentationLexical NormalizationPart-Of-Speech TaggingPOS+1To What Extent Does Lexical Normalization Help English-as-a-Second Language Learners to Read Noisy English Texts?
How difficult is it for English-as-a-second language (ESL) learners to read noisy English texts? Do ESL learners need lexical normalization to read noisy English texts? These questions may also affect community formation…
Lexical NormalizationMultilingual Sequence Labeling Approach to solve Lexical Normalization
The task of converting a nonstandard text to a standard and readable text is known as lexical normalization. Almost all the Natural Language Processing (NLP) applications require the text data in normalized form to build…
Language ModellingLexical NormalizationWord AlignmentSesame Street to Mount Sinai: BERT-constrained character-level Moses models for multilingual lexical normalization
This paper describes the HEL-LJU submissions to the MultiLexNorm shared task on multilingual lexical normalization. Our system is based on a BERT token classification preprocessing step, where for each token the type of …
Lexical Normalizationtoken-classificationToken ClassificationMultiLexNorm: A Shared Task on Multilingual Lexical Normalization
Lexical normalization is the task of transforming an utterance into its standardized form. This task is beneficial for downstream analysis, as it provides a way to harmonize (often spontaneous) linguistic variation. Such…
Dependency ParsingLexical NormalizationPart-Of-Speech TaggingPOSCL-MoNoise: Cross-lingual Lexical Normalization
Social media is notoriously difficult to process for existing natural language processing tools, because of spelling errors, non-standard words, shortenings, non-standard capitalization and punctuation. One method to cir…
Lexical NormalizationÚFAL at MultiLexNorm 2021: Improving Multilingual Lexical Normalization by Fine-tuning ByT5
We present the winning entry to the Multilingual Lexical Normalization (MultiLexNorm) shared task at W-NUT 2021 (van der Goot et al., 2021a), which evaluates lexical-normalization systems on 12 social media datasets in 1…
Dependency ParsingLanguage ModelingLanguage ModellingLexical NormalizationContrastive String Representation Learning using Synthetic Data
String representation Learning (SRL) is an important task in the field of Natural Language Processing, but it remains under-explored. The goal of SRL is to learn dense and low-dimensional vectors (or embeddings) for enco…
Contrastive LearningLexical NormalizationRepresentation LearningSequence-to-Sequence Lexical Normalization with Multilingual Transformers
Current benchmark tasks for natural language processing contain text that is qualitatively different from the text used in informal day to day digital communication. This discrepancy has led to severe performance degrada…
Lexical NormalizationMachine TranslationSentenceTranslationDaN+: Danish Nested Named Entities and Lexical Normalization
This paper introduces DaN+, a new multi-domain corpus and annotation guidelines for Danish nested named entities (NEs) and lexical normalization to support research on cross-lingual cross-domain learning for a less-resou…
Cross-Lingual TransferLexical NormalizationMulti-Task Learningnamed-entity-recognition+3User-Generated Text Corpus for Evaluating Japanese Morphological Analysis and Lexical Normalization
Morphological analysis (MA) and lexical normalization (LN) are both important tasks for Japanese user-generated text (UGT). To evaluate and compare different MA/LN systems, we have constructed a publicly available Japane…
Lexical NormalizationMorphological AnalysisLexical Normalization for Code-switched Data and its Effect on POS Tagging
Lexical normalization, the translation of non-canonical data to standard language, has shown to improve the performance of many natural language processing tasks on social media. Yet, using multiple languages in one utte…
Lexical NormalizationPOSPOS TaggingTranslation