NSL-MT: Linguistically Informed Negative Samples for Efficient Machine Translation in Low-Resource Languages
We introduce negative space learning machine translation (NSL-MT), a training method for underresourced languages, that augments limited parallel data with synthetically generated violations of the target language's grammar and explicitly penalizes the model when it assigns high probability to these linguistically invalid outputs. NSL-MT delivers improvements across all baselines we tested, including 3-12% BLEU gains for well-performing models and 56-89% gains for models lacking decent initial support. Furthermore, NSL-MT provides a 5x data efficiency multiplier: training with 1,000 examples matches or exceeds normal training with 5,000 examples. NSL-MT thus provides a data-efficient alternative training method for settings where parallel data is limited.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationSimilar Papers 제목 키워드 기반
Semantically-Informed Syntactic Machine Translation: A Tree-Grafting Approach
We describe a unified and coherent syntactic framework for supporting a semantically-informed syntactic approach to statistical machine translation. Semantically enriched syntactic tags assigned to the target-language tr…
Machine TranslationTranslationMorphological and Language-Agnostic Word Segmentation for NMT
The state of the art of handling rich morphology in neural machine translation (NMT) is to break word forms into subword units, so that the overall vocabulary size of these units fits the practical limits given by the NM…
GPUMachine TranslationNMTTranslationSign Language Machine Translation and the Sign Language Lexicon: A Linguistically Informed Approach
Natural language processing and the machine translation of spoken language (speech/text) has benefitted from significant scientific research and development in re-cent times, rapidly advancing the field. On the other han…
Machine TranslationTranslationFine-grained evaluation of Quality Estimation for Machine translation based on a linguistically-motivated Test Suite
We present an alternative method of evaluating Quality Estimation systems, which is based on a linguistically-motivated Test Suite. We create a test-set consisting of 14 linguistic error categories and we gather for each…
Machine TranslationTranslationBhashaSetu: A Data-Centric Approach to Low-Resource Machine Translation
We present BhashaSetu, a linguistically enriched English--Marathi parallel dataset addressing persistent data limitations in low-resource neural machine translation (NMT). Marathi, spoken by over 95 million people, remai…
Low-Resource Neural Machine Translationparameter-efficient fine-tuning