paper-with-me

홈 › Papers

Building the Macedonian-Croatian Parallel Corpus

2016-05-01 · LREC 2016 5 · Ines Cebovi{\'c}, Marko Tadi{\'c}

In this paper we present the newly created parallel corpus of two under-resourced languages, namely, Macedonian-Croatian Parallel Corpus (mk-hr{\_}pcorp) that has been collected during 2015 at the Faculty of Humanities and Social Sciences, University of Zagreb. The mk-hr{\_}pcorp is a unidirectional (mk→hr) parallel corpus composed of synchronic fictional prose texts received already in digital form with over 500 Kw in each language. The corpus was sentence segmented and provides 39,735 aligned sentences. The alignment was done automatically and then post-corrected manually. The alignments order was shuffled and this enabled the corpus to be available under CC-BY license through META-SHARE. However, this prevents the research in language units over the sentence level.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

Building the Spanish-Croatian Parallel Corpus

2020-05-01 · LREC 2020 5 · Bojana Mikeleni{\'c}, Marko Tadi{\'c}

This paper describes the building of the first Spanish-Croatian unidirectional parallel corpus, which has been constructed at the Faculty of Humanities and Social Sciences of the University of Zagreb. The corpus is compr…

POSSentenceTranslation

Multi-lingual Argumentative Corpora in English, Turkish, Greek, Albanian, Croatian, Serbian, Macedonian, Bulgarian, Romanian and Arabic

2018-05-01 · LREC 2018 5 · Alfred Sliwa, Yuan Ma, Ruishen Liu, Niravkumar Borad 외
Argument MiningDecision Making

MULTEXT-East

2020-03-31 · Tomaž Erjavec

MULTEXT-East language resources, a multilingual dataset for language engineering research, focused on the morphosyntactic level of linguistic description. The MULTEXT-East dataset includes the EAGLES-based morphosyntacti…

Sentence

Comparing two acquisition systems for automatically building an English---Croatian parallel corpus from multilingual websites

2014-05-01 · LREC 2014 5 · Miquel Espl{\`a}-Gomis, Filip Klubi{\v{c}}ka, Nikola Ljube{\v{s}}i{\'c}, Sergio Ortiz-Rojas 외

In this paper we compare two tools for automatically harvesting bitexts from multilingual websites: bitextor and ILSP-FC. We used both tools for crawling 21 multilingual websites from the tourism domain to build a domain…

Information RetrievalMachine TranslationNatural Language Inference

Quality Estimation for Synthetic Parallel Data Generation

2014-05-01 · LREC 2014 5 · Raphael Rubino, Antonio Toral, Nikola Ljube{\v{s}}i{\'c}, Gema Ram{\'\i}rez-S{\'a}nchez

This paper presents a novel approach for parallel data generation using machine translation and quality estimation. Our study focuses on pivot-based machine translation from English to Croatian through Slovene. We genera…

Machine TranslationSentenceTranslation