Pseudo-Labels Are All You Need
Automatically estimating the complexity of texts for readers has a variety of applications, such as recommending texts with an appropriate complexity level to language learners or supporting the evaluation of text simplification approaches. In this paper, we present our submission to the Text Complexity DE Challenge 2022, a regression task where the goal is to predict the complexity of a German sentence for German learners at level B. Our approach relies on more than 220,000 pseudo-labels created from the German Wikipedia and other corpora to train Transformer-based models, and refrains from any feature engineering or any additional, labeled data. We find that the pseudo-label-based approach gives impressive results yet requires little to no adjustment to the specific task and therefore could be easily adapted to other domains and tasks.
Code (0)
등록된 구현이 없습니다.
Tasks
AllFeature EngineeringPseudo LabelSentenceText SimplificationSimilar Papers 제목 키워드 기반
Unsupervised Clustering using Pseudo-semi-supervised Learning
In this paper, we propose a framework that leverages semi-supervised models to improve unsupervised clustering performance. To leverage semi-supervised models, we first need to automatically generate labels, called pseud…
ClusteringTowards Understanding GD with Hard and Conjugate Pseudo-labels for Test-Time Adaptation
We consider a setting that a model needs to adapt to a new domain under distribution shifts, given that only unlabeled test samples from the new domain are accessible at test time. A common idea in most of the related wo…
Binary ClassificationDomain AdaptationTest-time AdaptationPseudo-Label Noise Suppression Techniques for Semi-Supervised Semantic Segmentation
Semi-supervised learning (SSL) can reduce the need for large labelled datasets by incorporating unlabelled data into the training. This is particularly interesting for semantic segmentation, where labelling data is very …
Pose EstimationPseudo LabelPseudo Label FilteringSemantic Segmentation+1Dense Teacher: Dense Pseudo-Labels for Semi-supervised Object Detection
To date, the most powerful semi-supervised object detectors (SS-OD) are based on pseudo-boxes, which need a sequence of post-processing with fine-tuned hyper-parameters. In this work, we propose replacing the sparse pseu…
object-detectionObject DetectionPseudo LabelSemi-Supervised Object DetectionSegment Anything Model (SAM) Enhanced Pseudo Labels for Weakly Supervised Semantic Segmentation
Weakly supervised semantic segmentation (WSSS) aims to bypass the need for laborious pixel-level annotation by using only image-level annotation. Most existing methods rely on Class Activation Maps (CAM) to derive pixel-…
ObjectPseudo LabelSemantic SegmentationWeakly supervised Semantic Segmentation+1