paper-with-me

홈 › Papers

Parallel Corpus of Croatian-Italian Administrative Texts

2019-09-01 · RANLP 2019 9 · Marija Brkic Bakaric, Ivana Lalli Pacelat

Parallel corpora constitute a unique re-source for providing assistance to human translators. The selection and preparation of the parallel corpora also conditions the quality of the resulting MT engine. Since Croatian is a national language and Italian is officially recognized as a minority lan-guage in seven cities and twelve munici-palities of Istria County, a large amount of parallel texts is produced on a daily basis. However, there have been no attempts in using these texts for compiling a parallel corpus. A domain-specific sentence-aligned parallel Croatian-Italian corpus of administrative texts would be of high value in creating different language tools and resources. The aim of this paper is, therefore, to explore the value of parallel documents which are publicly available mostly in pdf format and to investigate the use of automatically-built dictionaries in corpus compilation. The effects that a document format and, consequently sentence splitting, and the dictionary input have on the sentence alignment process are manually evaluated.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

Building the Macedonian-Croatian Parallel Corpus

2016-05-01 · LREC 2016 5 · Ines Cebovi{\'c}, Marko Tadi{\'c}

In this paper we present the newly created parallel corpus of two under-resourced languages, namely, Macedonian-Croatian Parallel Corpus (mk-hr{\_}pcorp) that has been collected during 2015 at the Faculty of Humanities a…

Sentence

DiHuTra: a Parallel Corpus to Analyse Differences between Human Translations

2022-06-01 · LREC 2022 6 · Ekaterina Lapshinova-Koltunski, Maja Popović, Maarit Koponen

This project aimed to design a corpus of parallel human translations (HTs) of the same source texts by professionals and students. The resulting corpus consists of English news and reviews source texts, their translation…

Machine TranslationTranslation

Building the Spanish-Croatian Parallel Corpus

2020-05-01 · LREC 2020 5 · Bojana Mikeleni{\'c}, Marko Tadi{\'c}

This paper describes the building of the first Spanish-Croatian unidirectional parallel corpus, which has been constructed at the Faculty of Humanities and Social Sciences of the University of Zagreb. The corpus is compr…

POSSentenceTranslation

Comparing two acquisition systems for automatically building an English---Croatian parallel corpus from multilingual websites

2014-05-01 · LREC 2014 5 · Miquel Espl{\`a}-Gomis, Filip Klubi{\v{c}}ka, Nikola Ljube{\v{s}}i{\'c}, Sergio Ortiz-Rojas 외

In this paper we compare two tools for automatically harvesting bitexts from multilingual websites: bitextor and ILSP-FC. We used both tools for crawling 21 multilingual websites from the tourism domain to build a domain…

Information RetrievalMachine TranslationNatural Language Inference

Quality Estimation for Synthetic Parallel Data Generation

2014-05-01 · LREC 2014 5 · Raphael Rubino, Antonio Toral, Nikola Ljube{\v{s}}i{\'c}, Gema Ram{\'\i}rez-S{\'a}nchez

This paper presents a novel approach for parallel data generation using machine translation and quality estimation. Our study focuses on pivot-based machine translation from English to Croatian through Slovene. We genera…

Machine TranslationSentenceTranslation