paper-with-me

Papers

Peru is Multilingual, Its Machine Translation Should Be Too?

2021-06-01 · NAACL (AmericasNLP) 2021 6 · Arturo Oncevay

Peru is a multilingual country with a long history of contact between the indigenous languages and Spanish. Taking advantage of this context for machine translation is possible with multilingual approaches for learning both unsupervised subword segmentation and neural machine translation models. The study proposes the first multilingual translation models for four languages spoken in Peru: Aymara, Ashaninka, Quechua and Shipibo-Konibo, providing both many-to-Spanish and Spanish-to-many models and outperforming pairwise baselines in most of them. The task exploited a large English-Spanish dataset for pre-training, monolingual texts with tagged back-translation, and parallel corpora aligned with English. Finally, by fine-tuning the best models, we also assessed the out-of-domain capabilities in two evaluation datasets for Quechua and a new one for Shipibo-Konibo.

📄 PDF Abstract BibTeX

Code (1)

aoncevay/mt-peru 공식 구현

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Huqariq: A Multilingual Speech Corpus of Native Languages of Peru forSpeech Recognition

2022-06-01 · LREC 2022 6 · Rodolfo Zevallos, Luis Camacho, Nelsi Melgarejo

The Huqariq corpus is a multilingual collection of speech from native Peruvian languages. The transcribed corpus is intended for the research and development of speech technologies to preserve endangered languages in Per…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Identificationspeech-recognition+3

Huqariq: A Multilingual Speech Corpus of Native Languages of Peru for Speech Recognition

2022-07-12 · Rodolfo Zevallos, Luis Camacho, Nelsi Melgarejo

The Huqariq corpus is a multilingual collection of speech from native Peruvian languages. The transcribed corpus is intended for the research and development of speech technologies to preserve endangered languages in Per…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Identificationspeech-recognition+3

Multi30K: Multilingual English-German Image Descriptions

2016-05-02 · WS 2016 8 · Desmond Elliott, Stella Frank, Khalil Sima'an, Lucia Specia

We introduce the Multi30K dataset to stimulate multilingual multimodal research. Recent advances in image description have been demonstrated on English-language datasets almost exclusively, but image description should n…

Image DescriptionMachine TranslationMultimodal Machine TranslationTranslation

SuperUDF: Self-supervised UDF Estimation for Surface Reconstruction

2023-08-28 · Hui Tian, Chenyang Zhu, Yifei Shi, Kai Xu

Learning-based surface reconstruction based on unsigned distance functions (UDF) has many advantages such as handling open surfaces. We propose SuperUDF, a self-supervised UDF learning which exploits a learned geometry p…

Surface Reconstruction

Multilingual Neural Machine Translation with Language Clustering

2019-08-25 · IJCNLP 2019 11 · Xu Tan, Jiale Chen, Di He, Yingce Xia 외

Multilingual neural machine translation (NMT), which translates multiple languages using a single model, is of great practical importance due to its advantages in simplifying the training process, reducing online mainten…

ClusteringMachine TranslationNMTTranslation