Models in the Wild: On Corruption Robustness of NLP Systems
Natural Language Processing models lack a unified approach to robustness testing. In this paper we introduce WildNLP - a framework for testing model stability in a natural setting where text corruptions such as keyboard errors or misspelling occur. We compare robustness of models from 4 popular NLP tasks: Q&A, NLI, NER and Sentiment Analysis by testing their performance on aspects introduced in the framework. In particular, we focus on a comparison between recent state-of-the- art text representations and non-contextualized word embeddings. In order to improve robust- ness, we perform adversarial training on se- lected aspects and check its transferability to the improvement of models with various cor- ruption types. We find that the high perfor- mance of models does not ensure sufficient robustness, although modern embedding tech- niques help to improve it. We release cor- rupted datasets and code for WildNLP frame- work for the community.
Code (0)
등록된 구현이 없습니다.
Tasks
NERSentiment AnalysisWord EmbeddingsSimilar Papers 제목 키워드 기반
RobustGait: Robustness Analysis for Appearance Based Gait Recognition
Appearance-based gait recognition have achieved strong performance on controlled datasets, yet systematic evaluation of its robustness to real-world corruptions and silhouette variability remains lacking. We present Robu…
Knowledge DistillationGait RecognitionSpeech Robust Bench: A Robustness Benchmark For Speech Recognition
As Automatic Speech Recognition (ASR) models become ever more pervasive, it is important to ensure that they make reliable predictions under corruptions present in the physical and digital world. We propose Speech Robust…
Adversarial RobustnessAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Robust Speech Recognition+2Benchmarking Robustness of 3D Object Detection to Common Corruptions
3D object detection is an important task in autonomous driving to perceive the surroundings. Despite the excellent performance, the existing 3D detectors lack the robustness to real-world corruptions caused by advers…
3D Object DetectionAutonomous DrivingBenchmarkingobject-detection+1Benchmarking Popular Classification Models' Robustness to Random and Targeted Corruptions
Text classification models, especially neural networks based models, have reached very high accuracy on many popular benchmark datasets. Yet, such models when deployed in real world applications, tend to perform badly. T…
BenchmarkingClassificationGeneral Classificationtext-classification+1Improving the Robustness of 3D Human Pose Estimation: A Benchmark and Learning from Noisy Input
Despite the promising performance of current 3D human pose estimation techniques, understanding and enhancing their generalization on challenging in-the-wild videos remain an open problem. In this work, we focus on the r…
3D Human Pose EstimationData AugmentationPose Estimation