paper-with-me

홈 › Papers

Comparing two acquisition systems for automatically building an English---Croatian parallel corpus from multilingual websites

2014-05-01 · LREC 2014 5 · Miquel Espl{\`a}-Gomis, Filip Klubi{\v{c}}ka, Nikola Ljube{\v{s}}i{\'c}, Sergio Ortiz-Rojas, Vassilis Papavassiliou, Prokopis Prokopidis

In this paper we compare two tools for automatically harvesting bitexts from multilingual websites: bitextor and ILSP-FC. We used both tools for crawling 21 multilingual websites from the tourism domain to build a domain-specific English―Croatian parallel corpus. Different settings were tried for both tools and 10,662 unique document pairs were obtained. A sample of about 10{\%} of them was manually examined and the success rate was computed on the collection of pairs of documents detected by each setting. We compare the performance of the settings and the amount of different corpora detected by each setting. In addition, we describe the resource obtained, both by the settings and through the human evaluation, which has been released as a high-quality parallel corpus.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalMachine TranslationNatural Language Inference

Similar Papers 제목 키워드 기반

OdiEnCorp 2.0: Odia-English Parallel Corpus for Machine Translation

2020-05-01 · LREC 2020 5 · Shantipriya Parida, Satya Ranjan Dash, Ond{\v{r}}ej Bojar, Petr Motlicek 외

The preparation of parallel corpora is a challenging task, particularly for languages that suffer from under-representation in the digital world. In a multi-lingual country like India, the need for such parallel corpora …

Machine TranslationNMTOptical Character RecognitionOptical Character Recognition (OCR)+1

An overview of Portuguese WordNets

2016-01-01 · GWC 2016 1 · Valeria de Paiva, Livy Real, Hugo Gonçalo Oliveira, Alexandre Rademaker 외

Semantic relations between words are key to building systems that aim to understand and manipulate language. For English, the “de facto” standard for representing this kind of knowledge is Princeton’s WordNet. Here, we d…

Producing Monolingual and Parallel Web Corpora at the Same Time - SpiderLing and Bitextor's Love Affair

2016-05-01 · LREC 2016 5 · Nikola Ljube{\v{s}}i{\'c}, Miquel Espl{\`a}-Gomis, Antonio Toral, Sergio Ortiz Rojas 외

This paper presents an approach for building large monolingual corpora and, at the same time, extracting parallel data by crawling the top-level domain of a given language of interest. For gathering linguistically releva…

Comparing Sense Categorization Between English PropBank and English WordNet

2019-07-01 · GWC 2019 7 · Özge Bakay, Begüm Avar, Olcay Taner Yildiz

Given the fact that verbs play a crucial role in language comprehension, this paper presents a study which compares the verb senses in English PropBank with the ones in English WordNet through manual tagging. After analy…

Unsupervised Pidgin Text Generation By Pivoting English Data and Self-Training

2020-03-18 · Ernie Chang, David Ifeoluwa Adelani, Xiaoyu Shen, Vera Demberg

West African Pidgin English is a language that is significantly spoken in West Africa, consisting of at least 75 million speakers. Nevertheless, proper machine translation systems and relevant NLP datasets for pidgin Eng…

Data-to-Text GenerationMachine TranslationText GenerationTranslation