Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel Corpora
We conducted a detailed analysis on the quality of web-mined corpora for two low-resource languages (making three language pairs, English-Sinhala, English-Tamil and Sinhala-Tamil). We ranked each corpus according to a similarity measure and carried out an intrinsic and extrinsic evaluation on different portions of this ranked corpus. We show that there are significant quality differences between different portions of web-mined corpora and that the quality varies across languages and datasets. We also show that, for some web-mined datasets, Neural Machine Translation (NMT) models trained with their highest-ranked 25k portion can be on par with human-curated datasets.
Code (1)
Tasks
Machine TranslationNMTTranslationSimilar Papers 제목 키워드 기반
Most Apparent Distortion: Full-Reference Image Quality Assessment and the Role of Strategy
The mainstream approach to image quality assessment has centered around accurately modeling the single most relevant strategy employed by the human visual system (HVS) when judging image quality (e.g., detecting visible…
Full reference image quality assessmentFull-Reference Image Quality AssessmentImage Quality AssessmentInflation and income inequality: Does the level of income inequality matter?
In the recent times of global Covid pandemic, the Federal Reserve has raised the concerns of upsurges in prices. Given the complexity of interaction between inflation and inequality, we examine whether the impact of infl…
A Comparative Study of Quality and Content-Based Spatial Pooling Strategies in Image Quality Assessment
The process of quantifying image quality consists of engineering the quality features and pooling these features to obtain a value or a map. There has been a significant research interest in designing the quality feature…
Image Quality AssessmentSSIMThe impact of conformer quality on learned representations of molecular conformer ensembles
Training machine learning models to predict properties of molecular conformer ensembles is an increasingly popular strategy to accelerate the conformational analysis of drug-like small molecules, reactive organic substra…
Representation LearningDoes VLN Pretraining Work with Nonsensical or Irrelevant Instructions?
Data augmentation via back-translation is common when pretraining Vision-and-Language Navigation (VLN) models, even though the generated instructions are noisy. But: does that noise matter? We find that nonsensical or ir…
Data AugmentationTranslationVision and Language Navigation