paper-with-me

홈 › Papers

Enhancing Translation Accuracy of Large Language Models through Continual Pre-Training on Parallel Data

2024-07-03 · Minato Kondo, Takehito Utsuro, Masaaki Nagata

In this paper, we propose a two-phase training approach where pre-trained large language models are continually pre-trained on parallel data and then supervised fine-tuned with a small amount of high-quality parallel data. To investigate the effectiveness of our proposed approach, we conducted continual pre-training with a 3.8B-parameter model and parallel data across eight different formats. We evaluate these methods on thirteen test sets for Japanese-to-English and English-to-Japanese translation. The results demonstrate that when utilizing parallel data in continual pre-training, it is essential to alternate between source and target sentences. Additionally, we demonstrated that the translation accuracy improves only for translation directions where the order of source and target sentences aligns between continual pre-training data and inference. In addition, we demonstrate that the LLM-based translation model is more robust in translating spoken language and achieves higher accuracy with less training data compared to supervised encoder-decoder models. We also show that the highest accuracy is achieved when the data for continual pre-training consists of interleaved source and target sentences and when tags are added to the source sentences.

📄 PDF Abstract BibTeX arXiv:2407.03145

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderTranslation

Similar Papers 제목 키워드 기반

DUAL-REFLECT: Enhancing Large Language Models for Reflective Translation through Dual Learning Feedback Mechanisms

2024-06-11 · Andong Chen, Lianzhang Lou, Kehai Chen, Xuefeng Bai 외

Recently, large language models (LLMs) enhanced by self-reflection have achieved promising performance on machine translation. The key idea is guiding LLMs to generate translation with human-like feedback. However, exist…

Machine TranslationTranslation

LLaVA-SLT: Visual Language Tuning for Sign Language Translation

2024-12-21 · Han Liang, Chengyu Huang, Yuecheng Xu, Cheng Tang 외

In the realm of Sign Language Translation (SLT), reliance on costly gloss-annotated datasets has posed a significant barrier. Recent advancements in gloss-free SLT methods have shown promise, yet they often largely lag b…

Sign Language TranslationTranslation

Exploring Large Language Models for Translating Romanian Computational Problems into English

2025-01-09 · Adrian Marius Dumitran, Adrian-Catalin Badea, Stefan-Gabriel Muscalu, Angela-Liliana Dumitran 외

Recent studies have suggested that large language models (LLMs) underperform on mathematical and computer science tasks when these problems are translated from Romanian into English, compared to their original Romanian f…

Translation

Grounding Natural Language to SQL Translation with Data-Based Self-Explanations

2024-11-05 · Yuankai Fan, Tonghui Ren, Can Huang, Zhenying He 외

Natural Language Interfaces for Databases empower non-technical users to interact with data using natural language (NL). Advanced approaches, utilizing either neural sequence-to-sequence or more recent sophisticated larg…

Translation

Refining Translations with LLMs: A Constraint-Aware Iterative Prompting Approach

2024-11-13 · Shangfeng Chen, Xiayang Shi, Pu Li, YinLin Li 외

Large language models (LLMs) have demonstrated remarkable proficiency in machine translation (MT), even without specific training on the languages in question. However, translating rare words in low-resource or domain-sp…

Machine TranslationRAGRetrieval-augmented GenerationTranslation