paper-with-me

Papers

A multi-source approach for Breton–French hybrid machine translation

2020-11-01 · EAMT 2020 11 · Víctor M. Sánchez-Cartagena, Mikel L. Forcada, Felipe Sánchez-Martínez

Corpus-based approaches to machine translation (MT) have difficulties when the amount of parallel corpora to use for training is scarce, especially if the languages involved in the translation are highly inflected. This problem can be addressed from different perspectives, including data augmentation, transfer learning, and the use of additional resources, such as those used in rule-based MT. This paper focuses on the hybridisation of rule-based MT and neural MT for the Breton–French under-resourced language pair in an attempt to study to what extent the rule-based MT resources help improve the translation quality of the neural MT system for this particular under-resourced language pair. We combine both translation approaches in a multi-source neural MT architecture and find out that, even though the rule-based system has a low performance according to automatic evaluation metrics, using it leads to improved translation quality.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationMachine TranslationTransfer LearningTranslation

Similar Papers 제목 키워드 기반

A prototype dependency treebank for Breton

2018-05-01 · JEPTALNRECITAL 2018 5 · Francis M. Tyers, Vinit Ravishankar

This paper describes the development of the first syntactically-annotated corpus of Breton. The corpus is part of the Universal Dependencies project. In the paper we describe how the corpus was prepared, some Breton-spec…

Review on the Existing Language Resources for Languages of France

2016-05-01 · LREC 2016 5 · Thibault Grouas, Val{\'e}rie Mapelli, Quentin Samier

With the support of the DGLFLF, ELDA conducted an inventory of existing language resources for the regional languages of France. The main aim of this inventory was to assess the exploitability of the identified resources…

Cultural Vocal Bursts Intensity PredictionDiversityTranslation

Common Voice: A Massively-Multilingual Speech Corpus

2019-12-13 · LREC 2020 5 · Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty 외

The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Identificationspeech-recognition+3

Processing Mutations in Breton with Finite-State Transducers

2014-08-01 · WS 2014 8 · Thierry Poibeau

Book Review: Biomedical Natural Language Processing by Kevin Bretonnel Cohen and Dina Demner-Fushman

2017-04-01 · CL 2017 4 · Jin-Dong Kim
Named Entity Recognition (NER)Relation Extraction