Robust to Noise Models in Natural Language Processing Tasks
There are a lot of noise texts surrounding a person in modern life. The traditional approach is to use spelling correction, yet the existing solutions are far from perfect. We propose robust to noise word embeddings model, which outperforms existing commonly used models, like fasttext and word2vec in different tasks. In addition, we investigate the noise robustness of current models in different natural language processing tasks. We propose extensions for modern models in three downstream tasks, i.e. text classification, named entity recognition and aspect extraction, which shows improvement in noise robustness over existing solutions.
Code (2)
Tasks
Aspect Extractionnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Spelling Correctiontext-classificationText ClassificationWord EmbeddingsMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Towards a Better Understanding of Noise in Natural Language Processing
In this paper, we propose a definition and taxonomy of various types of non-standard textual content – generally referred to as “noise” – in Natural Language Processing (NLP). While data pre-processing is undoubtedly imp…
Contextual Text Denoising with Masked Language Model
Recently, with the help of deep learning models, significant advances have been made in different Natural Language Processing (NLP) tasks. Unfortunately, state-of-the-art models are vulnerable to noisy texts. We propose …
DenoisingLanguage ModelingLanguage ModellingmodelContextual Text Denoising with Masked Language Models
Recently, with the help of deep learning models, significant advances have been made in different Natural Language Processing (NLP) tasks. Unfortunately, state-of-the-art models are vulnerable to noisy texts. We propose …
DenoisingLanguage ModelingLanguage ModellingCausally Denoise Word Embeddings Using Half-Sibling Regression
Distributional representations of words, also known as word vectors, have become crucial for modern natural language processing tasks due to their wide applications. Recently, a growing body of word vector postprocessing…
Causal InferenceregressionSentiment AnalysisWord EmbeddingsRevisiting Noise in Natural Language Processing for Computational Social Science
Computational Social Science (CSS) is an emerging field driven by the unprecedented availability of human-generated content for researchers. This field, however, presents a unique set of challenges due to the nature of t…
Optical Character Recognition (OCR)