A Quality-based Active Sample Selection Strategy for Statistical Machine Translation
This paper presents a new active learning technique for machine translation based on quality estimation of automatically translated sentences. It uses an error-driven strategy, i.e., it assumes that the more errors an automatically translated sentence contains, the more informative it is for the translation system. Our approach is based on a quality estimation technique which involves a wider range of features of the source text, automatic translation, and machine translation system compared to previous work. In addition, we enhance the machine translation system training data with post-edited machine translations of the sentences selected, instead of simulating this using previously created reference translations. We found that re-training systems with additional post-edited data yields higher quality translations regardless of the selection strategy used. We relate this to the fact that post-editions tend to be closer to source sentences as compared to references, making the rule extraction process more reliable.
Code (0)
등록된 구현이 없습니다.
Tasks
Active LearningMachine TranslationSentenceSentiment AnalysisTranslationSimilar Papers 제목 키워드 기반
RADS: Reinforcement Learning-Based Sample Selection Improves Transfer Learning in Low-resource and Imbalanced Clinical Settings
A common strategy in transfer learning is few shot fine-tuning, but its success is highly dependent on the quality of samples selected as training examples. Active learning methods such as uncertainty sampling and divers…
Reinforcement LearningTransfer LearningActive LearningActive Learning for Event Extraction with Memory-based Loss Prediction Model
Event extraction (EE) plays an important role in many industrial application scenarios, and high-quality EE methods require a large amount of manual annotation data to train supervised learning models. However, the cost …
Active LearningEvent ExtractionHybrid Active Learning via Deep Clustering for Video Action Detection
In this work, we focus on reducing the annotation cost for video action detection which requires costly frame-wise dense annotations. We study a novel hybrid active learning (AL) strategy which performs efficient lab…
Action DetectionActive LearningClusteringDeep Clustering+3Active Learning for Object Detection with Non-Redundant Informative Sampling
Curating an informative and representative dataset is essential for enhancing the performance of 2D object detectors. We present a novel active learning sampling strategy that addresses both the informativeness and diver…
Active LearningDiversityimage-classificationImage Classification+4SFedCA: Credit Assignment-Based Active Client Selection Strategy for Spiking Federated Learning
Spiking federated learning is an emerging distributed learning paradigm that allows resource-constrained devices to train collaboratively at low power consumption without exchanging local data. It takes advantage of both…
Federated Learning