paper-with-me

홈 › Papers

Challenges of Building Domain-Specific Parallel Corpora from Public Administration Documents

2022-06-01 · LREC (BUCC) 2022 6 · Filip Klubička, Lorena Kasunić, Danijel Blazsetin, Petra Bago

PRINCIPLE was a Connecting Europe Facility (CEF)-funded project that focused on the identification, collection and processing of language resources (LRs) for four European under-resourced languages (Croatian, Icelandic, Irish and Norwegian) in order to improve translation quality of eTranslation, an online machine translation (MT) tool provided by the European Commission. The collected LRs were used for the development of neural MT engines in order to verify the quality of the resources. For all four languages, a total of 66 LRs were collected and made available on the ELRC-SHARE repository under various licenses. For Croatian, we have collected and published 20 LRs: 19 parallel corpora and 1 glossary. The majority of data is in the general domain (72 % of translation units), while the rest is in the eJustice (23 %), eHealth (3 %) and eProcurement (2 %) Digital Service Infrastructures (DSI) domains. The majority of the resources were for the Croatian-English language pair. The data was donated by six data contributors from the public as well as private sector. In this paper we present a subset of 13 Croatian LRs developed based on public administration documents, which are all made freely available, as well as challenges associated with the data collection, cleaning and processing.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Methods 이 논문이 사용한 방법론

Golden Queue Managers 설명 없음

Similar Papers 제목 키워드 기반

Building Machine Translation System for Software Product Descriptions Using Domain-specific Sub-corpora Extraction

2022-09-01 · AMTA 2022 9 · Pintu Lohar, Sinead Madden, Edmond O’Connor, Maja Popovic 외

Building Machine Translation systems for a specific domain requires a sufficiently large and good quality parallel corpus in that domain. However, this is a bit challenging task due to the lack of parallel data in many d…

Machine TranslationSentenceSentence EmbeddingSentence-Embedding+1

The AMARA Corpus: Building Parallel Language Resources for the Educational Domain

2014-05-01 · LREC 2014 5 · Ahmed Abdelali, Francisco Guzman, Hassan Sajjad, Stephan Vogel

This paper presents the AMARA corpus of on-line educational content: a new parallel corpus of educational video subtitles, multilingually aligned for 20 languages, i.e. 20 monolingual corpora and 190 parallel corpora. Th…

Machine TranslationTranslation

Building Comparable Corpora for Assessing Multi-Word Term Alignment

2022-06-01 · LREC 2022 6 · Omar Adjali, Emmanuel Morin, Pierre Zweigenbaum

Recent work has demonstrated the importance of dealing with Multi-Word Terms (MWTs) in several Natural Language Processing applications. In particular, MWTs pose serious challenges for alignment and machine translation s…

Machine Translation

Harvesting comparable corpora and mining them for equivalent bilingual sentences using statistical classification and analogy- based heuristics

2015-11-18 · Krzysztof Wołk, Emilia Rejmund, Krzysztof Marasek

Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our new methodologies for minin…

General ClassificationMachine TranslationRetrievalTranslation

Building and Modelling Multilingual Subjective Corpora

2014-05-01 · LREC 2014 5 · Motaz Saad, David Langlois, Kamel Sma{\"\i}li

Building multilingual opinionated models requires multilingual corpora annotated with opinion labels. Unfortunately, such kind of corpora are rare. We consider opinions in this work as subjective or objective. In this pa…

Language ModellingMachine TranslationOpinion MiningSentiment Analysis+1