paper-with-me

홈 › Papers

Building a multilingual parallel corpus for human users

2012-05-01 · LREC 2012 5 · Alex Rosen, R, Martin Vav{\v{r}}{\'\i}n

We present the architecture and the current state of InterCorp, a multilingual parallel corpus centered around Czech, intended primarily for human users and consisting of written texts with a focus on fiction. Following an outline of its recent development and a comparison with some other multilingual parallel corpora we give an overview of the data collection procedure that covers text selection criteria, data format, conversion, alignment, lemmatization and tagging. Finally, we show a sample query using the web-based search interface and discuss challenges and prospects of the project.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Lemmatization

Similar Papers 제목 키워드 기반

Comparing two acquisition systems for automatically building an English---Croatian parallel corpus from multilingual websites

2014-05-01 · LREC 2014 5 · Miquel Espl{\`a}-Gomis, Filip Klubi{\v{c}}ka, Nikola Ljube{\v{s}}i{\'c}, Sergio Ortiz-Rojas 외

In this paper we compare two tools for automatically harvesting bitexts from multilingual websites: bitextor and ILSP-FC. We used both tools for crawling 21 multilingual websites from the tourism domain to build a domain…

Information RetrievalMachine TranslationNatural Language Inference

KC4MT: A High-Quality Corpus for Multilingual Machine Translation

2022-06-01 · LREC 2022 6 · Vinh Van Nguyen, Ha Nguyen, Huong Thanh Le, Thai Phuong Nguyen 외

The multilingual parallel corpus is an important resource for many applications of natural language processing (NLP). For machine translation, the size and quality of the training corpus mainly affects the quality of the…

Machine TranslationSentenceTranslationVocal Bursts Intensity Prediction

Building The Sense-Tagged Multilingual Parallel Corpus

2014-05-01 · LREC 2014 5 · Shan Wang, Francis Bond

Sense-annotated parallel corpora play a crucial role in natural language processing. This paper introduces our progress in creating such a corpus for Asian languages using English as a pivot, which is the first such corp…

PMIndia -- A Collection of Parallel Corpora of Languages of India

2020-01-27 · Barry Haddow, Faheem Kirefu

Parallel text is required for building high-quality machine translation (MT) systems, as well as for other multilingual NLP applications. For many South Asian languages, such data is in short supply. In this paper, we de…

Machine TranslationMultilingual NLPNMTSentence+1

The AMARA Corpus: Building Parallel Language Resources for the Educational Domain

2014-05-01 · LREC 2014 5 · Ahmed Abdelali, Francisco Guzman, Hassan Sajjad, Stephan Vogel

This paper presents the AMARA corpus of on-line educational content: a new parallel corpus of educational video subtitles, multilingually aligned for 20 languages, i.e. 20 monolingual corpora and 190 parallel corpora. Th…

Machine TranslationTranslation