paper-with-me

홈 › Papers

Toward domain-specific machine translation and quality estimation systems

2026-03-26 · Javad Pourmostafa Roshan Sharami arxiv

Machine Translation (MT) and Quality Estimation (QE) perform well in general domains but degrade under domain mismatch. This dissertation studies how to adapt MT and QE systems to specialized domains through a set of data-focused contributions. Chapter 2 presents a similarity-based data selection method for MT. Small, targeted in-domain subsets outperform much larger generic datasets and reach strong translation quality at lower computational cost. Chapter 3 introduces a staged QE training pipeline that combines domain adaptation with lightweight data augmentation. The method improves performance across domains, languages, and resource settings, including zero-shot and cross-lingual cases. Chapter 4 studies the role of subword tokenization and vocabulary in fine-tuning. Aligned tokenization-vocabulary setups lead to stable training and better translation quality, while mismatched configurations reduce performance. Chapter 5 proposes a QE-guided in-context learning method for large language models. QE models select examples that improve translation quality without parameter updates and outperform standard retrieval methods. The approach also supports a reference-free setup, reducing reliance on a single reference set. These results show that domain adaptation depends on data selection, representation, and efficient adaptation strategies. The dissertation provides methods for building MT and QE systems that perform reliably in domain-specific settings.

📄 PDF Abstract BibTeX arXiv:2603.24955

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationData AugmentationDomain Adaptation

Similar Papers 제목 키워드 기반

Guiding In-Context Learning of LLMs through Quality Estimation for Machine Translation

2024-06-12 · Javad PourMostafa Roshan Sharami, Dimitar Shterionov, Pieter Spronck

The quality of output from large language models (LLMs), particularly in machine translation (MT), is closely tied to the quality of in-context examples (ICEs) provided along with the query, i.e., the text to translate. …

In-Context LearningLanguage ModelingLanguage ModellingMachine Translation+1

Domain-Specific Quality Estimation for Machine Translation in Low-Resource Scenarios

2026-03-07 · Namrata Patil Gurav, Akashdeep Ranu, Archchana Sindhujan, Diptesh Kanojia arxiv

Quality Estimation (QE) is essential for assessing machine translation quality in reference-less settings, particularly for domain-specific and low-resource language scenarios. In this paper, we investigate sentence-leve…

Machine Translation

Tailoring Domain Adaptation for Machine Translation Quality Estimation

2023-04-18 · Javad PourMostafa Roshan Sharami, Dimitar Shterionov, Frédéric Blain, Eva Vanmassenhove 외

While quality estimation (QE) can play an important role in the translation process, its effectiveness relies on the availability and quality of training data. For QE in particular, high-quality labeled data is often lac…

Data AugmentationDomain AdaptationMachine TranslationTranslation+1

Findings of the WMT 2018 Shared Task on Quality Estimation

2018-10-01 · WS 2018 10 · Lucia Specia, Fr{\'e}d{\'e}ric Blain, Varvara Logacheva, Ram{\'o}n Astudillo 외

We report the results of the WMT18 shared task on Quality Estimation, i.e. the task of predicting the quality of the output of machine translation systems at various granularity levels: word, phrase, sentence and documen…

Machine TranslationSentenceTranslation

Machine Translation Quality Estimation Across Domains

2014-08-01 · COLING 2014 8 · Jos{\'e} G. C. de Souza, Marco Turchi, Matteo Negri
Machine TranslationTranslation