paper-with-me

Papers

Corpus Creation and Initial SMT Experiments between Spanish and Shipibo-konibo

2017-09-01 · RANLP 2017 9 · Ana-Paula Galarreta, Andr{\'e}s Melgar, Arturo Oncevay

In this paper, we present the first attempts to develop a machine translation (MT) system between Spanish and Shipibo-konibo (es-shp). There are very few digital texts written in Shipibo-konibo and even less bilingual texts that can be aligned, hence we had to create a parallel corpus using both bilingual and monolingual texts. We will describe how this corpus was made, as well as the process we followed to improve the quality of the sentences used to build a statistical MT model or SMT. The results obtained surpassed the baseline proposed (dictionary based) and made a promising result for further development considering the size of corpus used. Finally, it is expected that this MT system can be reinforced with the use of additional linguistic rules and automatic language processing functions that are being implemented.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

An Annotated Corpus of Emerging Anglicisms in Spanish Newspaper Headlines

2020-04-06 · Elena Álvarez-Mellado

The extraction of anglicisms (lexical borrowings from English) is relevant both for lexicographic purposes and for NLP downstream tasks. We introduce a corpus of European Spanish newspaper headlines annotated with anglic…

An Annotated Corpus of Emerging Anglicisms in Spanish Newspaper Headlines

2020-05-01 · LREC 2020 5 · Elena Alvarez-Mellado

The extraction of anglicisms (lexical borrowings from English) is relevant both for lexicographic purposes and for NLP downstream tasks. We introduce a corpus of European Spanish newspaper headlines annotated with anglic…

Creation of an Annotated Corpus of Spanish Radiology Reports

2017-10-30 · Viviana Cotik, Darío Filippo, Roland Roller, Hans Uszkoreit 외

This paper presents a new annotated corpus of 513 anonymized radiology reports written in Spanish. Reports were manually annotated with entities, negation and uncertainty terms and relations. The corpus was conceived as …

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Negation+1

Automatic creation of WordNets from parallel corpora

2014-05-01 · LREC 2014 5 · Antoni Oliver, Salvador Climent

In this paper we present the evaluation results for the creation of WordNets for five languages (Spanish, French, German, Italian and Portuguese) using an approach based on parallel corpora. We have used three very large…

Information RetrievalTranslationWord Alignment

EPIC UdS - Creation and Applications of a Simultaneous Interpreting Corpus

2022-06-01 · LREC 2022 6 · Heike Przybyl, Ekaterina Lapshinova-Koltunski, Katrin Menzel, Stefan Fischer 외

In this paper, we describe the creation and annotation of EPIC UdS, a multilingual corpus of simultaneous interpreting for English, German and Spanish. We give an overview of the comparable and parallel, aligned corpus v…