paper-with-me

홈 › Papers

Building the Spanish-Croatian Parallel Corpus

2020-05-01 · LREC 2020 5 · Bojana Mikeleni{\'c}, Marko Tadi{\'c}

This paper describes the building of the first Spanish-Croatian unidirectional parallel corpus, which has been constructed at the Faculty of Humanities and Social Sciences of the University of Zagreb. The corpus is comprised of eleven Spanish novels and their translations to Croatian done by six different professional translators. All the texts were published between 1999 and 2012. The corpus has more than 2 Mw, with approximately 1 Mw for each language. It was automatically sentence segmented and aligned, as well as manually post-corrected, and contains 71,778 translation units. In order to protect the copyright and to make the corpus available under permissive CC-BY licence, the aligned translation units are shuffled. This limits the usability of the corpus for research of language units at sentence and lower language levels only. There are two versions of the corpus in TMX format that will be available for download through META-SHARE and CLARIN ERIC infrastructure. The former contains plain TMX, while the latter is lemmatised and POS-tagged and stored in the aTMX format.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

POSSentenceTranslation

Similar Papers 제목 키워드 기반

The EuroPat Corpus: A Parallel Corpus of European Patent Data

2022-06-01 · LREC 2022 6 · Kenneth Heafield, Elaine Farrow, Jelmer Van der Linde, Gema Ramírez-Sánchez 외

We present the EuroPat corpus of patent-specific parallel data for 6 official European languages paired with English: German, Spanish, French, Croatian, Norwegian, and Polish. The filtered parallel corpora range in size …

Machine TranslationTranslation

Building the Macedonian-Croatian Parallel Corpus

2016-05-01 · LREC 2016 5 · Ines Cebovi{\'c}, Marko Tadi{\'c}

In this paper we present the newly created parallel corpus of two under-resourced languages, namely, Macedonian-Croatian Parallel Corpus (mk-hr{\_}pcorp) that has been collected during 2015 at the Faculty of Humanities a…

Sentence

Comparing two acquisition systems for automatically building an English---Croatian parallel corpus from multilingual websites

2014-05-01 · LREC 2014 5 · Miquel Espl{\`a}-Gomis, Filip Klubi{\v{c}}ka, Nikola Ljube{\v{s}}i{\'c}, Sergio Ortiz-Rojas 외

In this paper we compare two tools for automatically harvesting bitexts from multilingual websites: bitextor and ILSP-FC. We used both tools for crawling 21 multilingual websites from the tourism domain to build a domain…

Information RetrievalMachine TranslationNatural Language Inference

Quality Estimation for Synthetic Parallel Data Generation

2014-05-01 · LREC 2014 5 · Raphael Rubino, Antonio Toral, Nikola Ljube{\v{s}}i{\'c}, Gema Ram{\'\i}rez-S{\'a}nchez

This paper presents a novel approach for parallel data generation using machine translation and quality estimation. Our study focuses on pivot-based machine translation from English to Croatian through Slovene. We genera…

Machine TranslationSentenceTranslation

ReSiPC: a Tool for Complex Searches in Parallel Corpora

2020-05-01 · LREC 2020 5 · Antoni Oliver, Bojana Mikeleni{\'c}

In this paper, a tool specifically designed to allow for complex searches in large parallel corpora is presented. The formalism for the queries is very powerful as it uses standard regular expressions that allow for comp…

POSTAG