paper-with-me

홈 › Papers

Building a Corpus for Corporate Websites Machine Translation Evaluation. A Step by Step Methodological Approach

2021-07-01 · TRITON 2021 7 · Irene Rivera-Trigueros, María-Dolores Olvera-Lobo

The aim of this paper is to describe the process carried out to develop a paral-lel corpus comprised of texts extracted from the corporate websites of south-ern Spanish SMEs from the sanitary sector which will serve as the basis for MT quality assessment. The stages for compiling the parallel corpora were: (i) selection of websites with content translated in English and Spanish, (ii) downloading of the HTML files of the selected websites, (iii) files filtering and pairing of English files with their Spanish equivalents, (iv) compilation of individual corpora (EN and ES) for each of the selected websites, (v) merging of the individual corpora into a two general corpus one in English and the other in Spanish, (vi) selection a representative sample of segments to be used as original (ES) and reference translations (EN), (vii) building of the parallel corpus intended for MT evaluation. The parallel corpus generated will serve to future Machine Translation quality assessment. In addition, the monolingual corpora generated during the process could as a base to carry out research focused on linguistic – bilingual or monolingual − analysis.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

Leveraging Multilingual News Websites for Building a Kurdish Parallel Corpus

2020-10-04 · Sina Ahmadi, Hossein Hassani, Daban Q. Jaff

Machine translation has been a major motivation of development in natural language processing. Despite the burgeoning achievements in creating more efficient machine translation systems thanks to deep learning methods, p…

ArticlesMachine TranslationTranslationTransliteration

OdiEnCorp 2.0: Odia-English Parallel Corpus for Machine Translation

2020-05-01 · LREC 2020 5 · Shantipriya Parida, Satya Ranjan Dash, Ond{\v{r}}ej Bojar, Petr Motlicek 외

The preparation of parallel corpora is a challenging task, particularly for languages that suffer from under-representation in the digital world. In a multi-lingual country like India, the need for such parallel corpora …

Machine TranslationNMTOptical Character RecognitionOptical Character Recognition (OCR)+1

Chinese-Portuguese Machine Translation: A Study on Building Parallel Corpora from Comparable Texts

2018-04-05 · LREC 2018 5 · Siyou Liu, Long-Yue Wang, Chao-Hong Liu

Although there are increasing and significant ties between China and Portuguese-speaking countries, there is not much parallel corpora in the Chinese-Portuguese language pair. Both languages are very populous, with 1.2 b…

Machine TranslationNMTTranslation

Revisiting Low Resource Status of Indian Languages in Machine Translation

2020-08-11 · Jerin Philip, Shashank Siripragada, Vinay P. Namboodiri, C. V. Jawahar

Indian language machine translation performance is hampered due to the lack of large scale multi-lingual sentence aligned corpora and robust benchmarks. Through this paper, we provide and analyse an automated framework t…

Machine TranslationNMTRetrievalSentence+1

Evaluating Machine Translation Models for English-Hindi Language Pairs: A Comparative Analysis

2025-05-26 · Ahan Prasannakumar Shetty

Machine translation has become a critical tool in bridging linguistic gaps, especially between languages as diverse as English and Hindi. This paper comprehensively evaluates various machine translation models for transl…

Machine TranslationTranslation