Learning-From-Mistakes Prompting for Indigenous Language Translation
Using large language models, this paper presents techniques to improve extremely low-resourced indigenous language translations. Our approaches are grounded in the use of (1) the presence of a datastore consisting of a limited number of parallel translation examples, (2) the inherent capabilities of LLMs like GPT-3.5, and (3) a word-level translation dictionary. We harness the potential of LLMs and in-context learning techniques in such a setting for using LLMs as universal translators for extremely low-resourced languages. Our methodology hinges on utilizing LLMs as language compilers for selected language pairs, hypothesizing that they could internalize syntactic structures to facilitate accurate translation. We introduce three techniques: KNNPrompting with Retrieved Prompting Context, Chain-of-Thought Prompting and Learningfrom-Mistakes Prompting, with the last method addressing past errors. The evaluation results suggest that, even with limited corpora, LLMs can effectively translate extremely low-resource languages when paired with proper prompting.
Code (0)
등록된 구현이 없습니다.
Tasks
In-Context LearningTranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
IndT5: A Text-to-Text Transformer for 10 Indigenous Languages
Transformer language models have become fundamental components of natural language processing based pipelines. Although several Transformer models have been introduced to serve many languages, there is a shortage of mode…
Language ModelingLanguage ModellingMachine TranslationTranslationEthical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers
In recent years machine translation has become very successful for high-resource language pairs. This has also sparked new interest in research on the automatic translation of low-resource languages, including Indigenous…
Machine TranslationTranslationParallel Corpus for Indigenous Language Translation: Spanish-Mazatec and Spanish-Mixtec
In this paper, we present a parallel Spanish-Mazatec and Spanish-Mixtec corpus for machine translation (MT) tasks, where Mazatec and Mixtec are two indigenous Mexican languages. We evaluated the usability of the collecte…
Few-Shot LearningMachine TranslationTransfer LearningTranslationRevitalization of Indigenous Languages through Pre-processing and Neural Machine Translation: The case of Inuktitut
Indigenous languages have been very challenging when dealing with NLP tasks and applications because of multiple reasons. These languages, in linguistic typology, are polysynthetic and highly inflected with rich morphoph…
Machine TranslationMorphological AnalysisTranslationSheffield's Submission to the AmericasNLP Shared Task on Machine Translation into Indigenous Languages
In this paper we describe the University of Sheffield's submission to the AmericasNLP 2023 Shared Task on Machine Translation into Indigenous Languages which comprises the translation from Spanish to eleven indigenous la…
AllArticlesMachine TranslationTranslation