Papers Lexical Normalization
“Lexical Normalization” 태그가 달린 논문 49편 · 필터 해제
Lexical Normalization for Code-switched Data and its Effect on POS-tagging
Lexical normalization, the translation of non-canonical data to standard language, has shown to improve the performance of manynatural language processing tasks on social media. Yet, using multiple languages in one utter…
Language IdentificationLexical NormalizationPOSPOS Tagging+1Norm It! Lexical Normalization for Italian and Its Downstream Effects for Dependency Parsing
Lexical normalization is the task of translating non-standard social media data to a standard form. Previous work has shown that this is beneficial for many downstream tasks in multiple languages. However, for Italian, t…
Dependency ParsingLexical NormalizationSynthetic Data for English Lexical Normalization: How Close Can We Get to Manually Annotated Data?
Social media is a valuable data resource for various natural language processing (NLP) tasks. However, standard NLP tools were often designed with standard texts in mind, and their performance decreases heavily when appl…
Lexical NormalizationSentenceWord EmbeddingsA Clustering Framework for Lexical Normalization of Roman Urdu
Roman Urdu is an informal form of the Urdu language written in Roman script, which is widely used in South Asia for online textual content. It lacks standard spelling and hence poses several normalization challenges duri…
ClusteringLexical NormalizationAdapting Deep Learning for Sentiment Classification of Code-Switched Informal Short Text
Nowadays, an abundance of short text is being generated that uses nonstandard writing styles influenced by regional languages. Such informal and code-switched content are under-resourced in terms of labeled datasets and …
ClassificationGeneral ClassificationLexical NormalizationSentiment Analysis+2A Multi-cascaded Deep Model for Bilingual SMS Classification
Most studies on text classification are focused on the English language. However, short texts such as SMS are influenced by regional languages. This makes the automatic text classification task challenging due to the mul…
ClassificationGeneral ClassificationLexical NormalizationMultilingual text classification+4An In-depth Analysis of the Effect of Lexical Normalization on the Dependency Parsing of Social Media
Existing natural language processing systems have often been designed with standard texts in mind. However, when these tools are used on the substantially different texts from social media, their performance drops dramat…
Dependency ParsingLexical NormalizationEnhancing BERT for Lexical Normalization
Language model-based pre-trained representations have become ubiquitous in natural language processing. They have been shown to significantly improve the performance of neural models on a great variety of tasks. However,…
Language ModelingLanguage ModellingLexical NormalizationNormalization of Indonesian-English Code-Mixed Twitter Data
Twitter is an excellent source of data for NLP researches as it offers tremendous amount of textual data. However, processing tweet to extract meaningful information is very challenging, at least for two reasons: (i) usi…
Language IdentificationLexical NormalizationTranslationLexical Normalization of User-Generated Medical Text
In the medical domain, user-generated social media text is increasingly used as a valuable complementary knowledge source to scientific medical literature. The extraction of this knowledge is complicated by colloquial la…
Lexical NormalizationMistake DetectionSpelling CorrectionMoNoise: A Multi-lingual and Easy-to-use Lexical Normalization Tool
In this paper, we introduce and demonstrate the online demo as well as the command line interface of a lexical normalization system (MoNoise) for a variety of languages. We further improve this model by using features fr…
Lexical NormalizationAdapting Sequence to Sequence models for Text Normalization in Social Media
Social media offer an abundant source of valuable raw data, however informal writing can quickly become a bottleneck for many natural language processing (NLP) tasks. Off-the-shelf tools are usually trained on formal tex…
DecoderLexical NormalizationText NormalizationModeling Input Uncertainty in Neural Network Dependency Parsing
Recently introduced neural network parsers allow for new approaches to circumvent data sparsity issues by modeling character level information and by exploiting raw data in a semi-supervised setting. Data sparsity is esp…
Dependency ParsingLexical NormalizationWord EmbeddingsNoise-Robust Morphological Disambiguation for Dialectal Arabic
User-generated text tends to be noisy with many lexical and orthographic inconsistencies, making natural language processing (NLP) tasks more challenging. The challenging nature of noisy text processing is exacerbated fo…
Lexical NormalizationMorphological AnalysisMorphological DisambiguationMorphological Tagging+1Handling Normalization Issues for Part-of-Speech Tagging of Online Conversational Text
A Taxonomy for In-depth Evaluation of Normalization for User Generated Content
MoNoise: Modeling Noise Using a Modular Normalization System
We propose MoNoise: a normalization model focused on generalizability and efficiency, it aims at being easily reusable and adaptable. Normalization is the task of translating texts from a non- canonical domain to a more …
Lexical NormalizationSpelling CorrectionWord EmbeddingsThe Denoised Web Treebank: Evaluating Dependency Parsing under Noisy Input Conditions
We introduce the Denoised Web Treebank: a treebank including a normalization layer and a corresponding evaluation metric for dependency parsing of noisy text, such as Tweets. This benchmark enables the evaluation of pars…
Dependency ParsingLexical NormalizationMachine TranslationText Normalization+1