Comparing two acquisition systems for automatically building an English---Croatian parallel corpus from multilingual websites
In this paper we compare two tools for automatically harvesting bitexts from multilingual websites: bitextor and ILSP-FC. We used both tools for crawling 21 multilingual websites from the tourism domain to build a domain-specific English―Croatian parallel corpus. Different settings were tried for both tools and 10,662 unique document pairs were obtained. A sample of about 10{\%} of them was manually examined and the success rate was computed on the collection of pairs of documents detected by each setting. We compare the performance of the settings and the amount of different corpora detected by each setting. In addition, we describe the resource obtained, both by the settings and through the human evaluation, which has been released as a high-quality parallel corpus.
Code (0)
등록된 구현이 없습니다.
Tasks
Information RetrievalMachine TranslationNatural Language InferenceSimilar Papers 제목 키워드 기반
OdiEnCorp 2.0: Odia-English Parallel Corpus for Machine Translation
The preparation of parallel corpora is a challenging task, particularly for languages that suffer from under-representation in the digital world. In a multi-lingual country like India, the need for such parallel corpora …
Machine TranslationNMTOptical Character RecognitionOptical Character Recognition (OCR)+1An overview of Portuguese WordNets
Semantic relations between words are key to building systems that aim to understand and manipulate language. For English, the “de facto” standard for representing this kind of knowledge is Princeton’s WordNet. Here, we d…
Producing Monolingual and Parallel Web Corpora at the Same Time - SpiderLing and Bitextor's Love Affair
This paper presents an approach for building large monolingual corpora and, at the same time, extracting parallel data by crawling the top-level domain of a given language of interest. For gathering linguistically releva…
Comparing Sense Categorization Between English PropBank and English WordNet
Given the fact that verbs play a crucial role in language comprehension, this paper presents a study which compares the verb senses in English PropBank with the ones in English WordNet through manual tagging. After analy…
Unsupervised Pidgin Text Generation By Pivoting English Data and Self-Training
West African Pidgin English is a language that is significantly spoken in West Africa, consisting of at least 75 million speakers. Nevertheless, proper machine translation systems and relevant NLP datasets for pidgin Eng…
Data-to-Text GenerationMachine TranslationText GenerationTranslation