paper-with-me

홈 › Papers

Design and compilation of a specialized Spanish-German parallel corpus

2012-05-01 · LREC 2012 5 · Carla Parra Escart{\'\i}n

This paper discusses the design and compilation of the TRIS corpus, a specialized parallel corpus of Spanish and German texts. It will be used for phraseological research aimed at improving statistical machine translation. The corpus is based on the European database of Technical Regulations Information System (TRIS), containing 995 original documents written in German and Spanish and their translations into Spanish and German respectively. This parallel corpus is under development and the first version with 97 aligned file pairs was released in the first META-NORD upload of metadata and resources in November 2011. The second version of the corpus, described in the current paper, contains 205 file pairs which have been completely aligned at sentence level, which account for approximately 1,563,000 words and 70,648 aligned sentence pairs.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentenceTranslation

Similar Papers 제목 키워드 기반

The EuroPat Corpus: A Parallel Corpus of European Patent Data

2022-06-01 · LREC 2022 6 · Kenneth Heafield, Elaine Farrow, Jelmer Van der Linde, Gema Ramírez-Sánchez 외

We present the EuroPat corpus of patent-specific parallel data for 6 official European languages paired with English: German, Spanish, French, Croatian, Norwegian, and Polish. The filtered parallel corpora range in size …

Machine TranslationTranslation

A Parallel Corpus Mixtec-Spanish

2019-08-01 · WS 2019 8 · Cynthia Monta{\~n}o, Gerardo Sierra Mart{\'\i}nez, Gemma Bel-Enguix, Helena Gomez

This work is about the compilation process of parallel documents Spanish-Mixtec. There are not many Spanish-Mixec parallel texts and most of the sources are non-digital books. Due to this, we need to face the errors when…

Sentence

Building a Corpus for Corporate Websites Machine Translation Evaluation. A Step by Step Methodological Approach

2021-07-01 · TRITON 2021 7 · Irene Rivera-Trigueros, María-Dolores Olvera-Lobo

The aim of this paper is to describe the process carried out to develop a paral-lel corpus comprised of texts extracted from the corporate websites of south-ern Spanish SMEs from the sanitary sector which will serve as t…

Machine Translation

4FX: Light Verb Constructions in a Multilingual Parallel Corpus

2014-05-01 · LREC 2014 5 · Anita R{\'a}cz, Istv{\'a}n Nagy T., Veronika Vincze

In this paper, we describe 4FX, a quadrilingual (English-Spanish-German-Hungarian) parallel corpus annotated for light verb constructions. We present the annotation process, and report statistical data on the frequency o…

Machine Translation

EMPAC: an English--Spanish Corpus of Institutional Subtitles

2020-05-01 · LREC 2020 5 · Iris Serrat Roozen, Jos{\'e} Manuel Mart{\'\i}nez Mart{\'\i}nez

The EuroparlTV Multimedia Parallel Corpus (EMPAC) is a collection of subtitles in English and Spanish for videos from the EuropeanParliament{'}s Multimedia Centre. The corpus has been compiled with the EMPAC toolkit. The…