Building Machine Translation System for Software Product Descriptions Using Domain-specific Sub-corpora Extraction
Building Machine Translation systems for a specific domain requires a sufficiently large and good quality parallel corpus in that domain. However, this is a bit challenging task due to the lack of parallel data in many domains such as economics, science and technology, sports etc. In this work, we build English-to-French translation systems for software product descriptions scraped from LinkedIn website. Moreover, we developed a first-ever test parallel data set of product descriptions. We conduct experiments by building a baseline translation system trained on general domain and then domain-adapted systems using sentence-embedding based corpus filtering and domain-specific sub-corpora extraction. All the systems are tested on our newly developed data set mentioned earlier. Our experimental evaluation reveals that the domain-adapted model based on our proposed approaches outperforms the baseline.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationSentenceSentence EmbeddingSentence-EmbeddingTranslationSimilar Papers 제목 키워드 기반
Automatic Standardization of Arabic Dialects for Machine Translation
Based on an annotated multimedia corpus, television series Mar{\=a}y{\=a} 2013, we dig into the question of ''automatic standardization'' of Arabic dialects for machine translation. Here we distinguish between rule-based…
Machine TranslationTranslationA Meta-Summary of Challenges in Building Products with ML Components -- Collecting Experiences from 4758+ Practitioners
Incorporating machine learning (ML) components into software products raises new software-engineering challenges and exacerbates existing challenges. Many researchers have invested significant effort in understanding the…
Collaboration Challenges in Building ML-Enabled Systems: Communication, Documentation, Engineering, and Process
The introduction of machine learning (ML) components in software projects has created the need for software engineers to collaborate with data scientists and other specialists. While collaboration can always be challengi…
FairnessToward a full-scale neural machine translation in production: the Booking.com use case
While some remarkable progress has been made in neural machine translation (NMT) research, there have not been many reports on its development and evaluation in practice. This paper tries to fill this gap by presenting s…
Machine TranslationNMTTranslationDoes Link Prediction Help Detect Feature Interactions in Software Product Lines (SPLs)?
An ongoing challenge for the requirements engineering of software product lines is to predict whether a new combination of features (units of functionality) will create an unwanted or even hazardous feature interaction. …
BIG-bench Machine LearningLink PredictionPrediction