paper-with-me

홈 › Papers

There's no Data Like Better Data: Using QE Metrics for MT Data Filtering

2023-11-09 · Jan-Thorsten Peter, David Vilar, Daniel Deutsch, Mara Finkelstein, Juraj Juraska, Markus Freitag

Quality Estimation (QE), the evaluation of machine translation output without the need of explicit references, has seen big improvements in the last years with the use of neural metrics. In this paper we analyze the viability of using QE metrics for filtering out bad quality sentence pairs in the training data of neural machine translation systems~(NMT). While most corpus filtering methods are focused on detecting noisy examples in collections of texts, usually huge amounts of web crawled data, QE models are trained to discriminate more fine-grained quality differences. We show that by selecting the highest quality sentence pairs in the training data, we can improve translation quality while reducing the training size by half. We also provide a detailed analysis of the filtering results, which highlights the differences between both approaches.

📄 PDF Abstract BibTeX arXiv:2311.05350

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationNMTSentenceTranslation

Similar Papers 제목 키워드 기반

What is the Best Automated Metric for Text to Motion Generation?

2023-09-19 · Jordan Voas, Yili Wang, QiXing Huang, Raymond Mooney

There is growing interest in generating skeleton-based human motions from natural language descriptions. While most efforts have focused on developing better neural architectures for this task, there has been no signific…

Motion Generation

Recall is the Proper Evaluation Metric for Word Segmentation

2017-11-01 · IJCNLP 2017 11 · Yan Shao, Christian Hardmeier, Joakim Nivre

We extensively analyse the correlations and drawbacks of conventionally employed evaluation metrics for word segmentation. Unlike in standard information retrieval, precision favours under-splitting systems and therefore…

Information RetrievalMachine TranslationPart-Of-Speech TaggingRetrieval+1

Human Evaluation and Correlation with Automatic Metrics in Consultation Note Generation

2022-04-01 · ACL 2022 5 · Francesco Moramarco, Alex Papadopoulos Korfiatis, Mark Perera, Damir Juric 외

In recent years, machine learning models have rapidly become better at generating clinical consultation notes; yet, there is little work on how to properly evaluate the generated consultation notes to understand the impa…

Injecting Planning-Awareness into Prediction and Detection Evaluation

2021-10-07 · Boris Ivanovic, Marco Pavone

Detecting other agents and forecasting their behavior is an integral part of the modern robotic autonomy stack, especially in safety-critical scenarios entailing human-robot interaction such as autonomous driving. Due to…

Autonomous DrivingDecision MakingPredictionTrajectory Forecasting

Is Preference Alignment Always the Best Option to Enhance LLM-Based Translation? An Empirical Analysis

2024-09-30 · Hippolyte Gisserot-Boukhlef, Ricardo Rei, Emmanuel Malherbe, Céline Hudelot 외

Neural metrics for machine translation (MT) evaluation have become increasingly prominent due to their superior correlation with human judgments compared to traditional lexical metrics. Researchers have therefore utilize…

Machine TranslationTranslation