paper-with-me

홈 › Papers

Domain Adaptation in MT Using Titles in Wikipedia as a Parallel Corpus: Resources and Evaluation

2016-05-01 · LREC 2016 5 · Gorka Labaka, I{\~n}aki Alegria, Kepa Sarasola

This paper presents how an state-of-the-art SMT system is enriched by using an extra in-domain parallel corpora extracted from Wikipedia. We collect corpora from parallel titles and from parallel fragments in comparable articles from Wikipedia. We carried out an evaluation with a double objective: evaluating the quality of the extracted data and evaluating the improvement due to the domain-adaptation. We think this can be very useful for languages with limited amount of parallel corpora, where in-domain data is crucial to improve the performance of MT sytems. The experiments on the Spanish-English language pair improve a baseline trained with the Europarl corpus in more than 2 points of BLEU when translating in the Computer Science domain.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesDomain Adaptation

Similar Papers 제목 키워드 기반

EduMT: Developing Machine Translation System for Educational Content in Indian Languages

2021-12-01 · ICON 2021 12 · Ramakrishna Appicharla, Asif Ekbal, Pushpak Bhattacharyya

In this paper, we explore various approaches to build Hindi to Bengali Neural Machine Translation (NMT) systems for the educational domain. Translation of educational content poses several challenges, such as unavailabil…

Data AugmentationDomain AdaptationMachine TranslationNMT+1

The AMARA Corpus: Building Parallel Language Resources for the Educational Domain

2014-05-01 · LREC 2014 5 · Ahmed Abdelali, Francisco Guzman, Hassan Sajjad, Stephan Vogel

This paper presents the AMARA corpus of on-line educational content: a new parallel corpus of educational video subtitles, multilingually aligned for 20 languages, i.e. 20 monolingual corpora and 190 parallel corpora. Th…

Machine TranslationTranslation

Generating Multilingual Parallel Corpus Using Subtitles

2018-04-11 · Farshad Jafari

Neural Machine Translation with its significant results, still has a great problem: lack or absence of parallel corpus for many languages. This article suggests a method for generating considerable amount of parallel cor…

Machine TranslationSentence

Models and Datasets for Cross-Lingual Summarisation

2022-02-19 · EMNLP 2021 11 · Laura Perez-Beltrachini, Mirella Lapata

We present a cross-lingual summarisation corpus with long documents in a source language associated with multi-sentence summaries in a target language. The corpus covers twelve language pairs and directions for four Euro…

ArticlesSentence

Assessing Gender Bias in Wikipedia: Inequalities in Article Titles

2021-08-01 · ACL (GeBNLP) 2021 8 · Agnieszka Falenska, Özlem Çetinoğlu

Potential gender biases existing in Wikipedia’s content can contribute to biased behaviors in a variety of downstream NLP systems. Yet, efforts in understanding what inequalities in portraying women and men occur in Wiki…

Articles