paper-with-me

홈 › Papers

The Joy of Parallelism with CzEng 1.0

2012-05-01 · LREC 2012 5 · Ond{\v{r}}ej Bojar, Zden{\v{e}}k {\v{Z}}abokrtsk{\'y}, Ond{\v{r}}ej Du{\v{s}}ek, Petra Galu{\v{s}}{\v{c}}{\'a}kov{\'a}, Martin Majli{\v{s}}, David Mare{\v{c}}ek, Ji{\v{r}}{\'\i} Mar{\v{s}}{\'\i}k, Michal Nov{\'a}k, Martin Popel, Ale{\v{s}} Tamchyna

CzEng 1.0 is an updated release of our Czech-English parallel corpus, freely available for non-commercial research or educational purposes. In this release, we approximately doubled the corpus size, reaching 15 million sentence pairs (about 200 million tokens per language). More importantly, we carefully filtered the data to reduce the amount of non-matching sentence pairs. CzEng 1.0 is automatically aligned at the level of sentences as well as words. We provide not only the plain text representation, but also automatic morphological tags, surface syntactic as well as deep syntactic dependency parse trees and automatic co-reference links in both English and Czech. This paper describes key properties of the released resource including the distribution of text domains, the corpus data formats, and a toolkit to handle the provided rich annotation. We also summarize the procedure of the rich annotation (incl. co-reference resolution) and of the automatic filtering. Finally, we provide some suggestions on exploiting such an automatically annotated sentence-parallel corpus.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentence

Similar Papers 제목 키워드 기반

Announcing CzEng 2.0 Parallel Corpus with over 2 Gigawords

2020-07-06 · Tom Kocmi, Martin Popel, Ondrej Bojar

We present a new release of the Czech-English parallel corpus CzEng 2.0 consisting of over 2 billion words (2 "gigawords") in each language. The corpus contains document-level information and is filtered with several tec…

Advancing Seq2seq with Joint Paraphrase Learning

2020-11-01 · EMNLP (ClinicalNLP) 2020 11 · So Yeon Min, Preethi Raghavan, Peter Szolovits

We address the problem of model generalization for sequence to sequence (seq2seq) architectures. We propose going beyond data augmentation via paraphrase-optimized multi-task learning and observe that it is useful in cor…

Data AugmentationMulti-Task LearningSemantic ParsingTranslation

Synonymy in Bilingual Context: The CzEngClass Lexicon

2018-08-01 · COLING 2018 8 · Zde{\v{n}}ka Ure{\v{s}}ov{\'a}, Eva Fu{\v{c}}{\'\i}kov{\'a}, Eva Haji{\v{c}}ov{\'a}, Jan Haji{\v{c}}

This paper describes CzEngClass, a bilingual lexical resource being built to investigate verbal synonymy in bilingual context and to relate semantic roles common to one synonym class to verb arguments (verb valency). In …

Word Sense Disambiguation

Learning to Identify Sentence Parallelism in Student Essays

2016-12-01 · COLING 2016 12 · Wei Song, Tong Liu, Ruiji Fu, Lizhen Liu 외

Parallelism is an important rhetorical device. We propose a machine learning approach for automated sentence parallelism identification in student essays. We build an essay dataset with sentence level parallelism annotat…

SentenceWord Alignment

Sequence Parallelism: Long Sequence Training from System Perspective

2021-05-26 · Shenggui Li, Fuzhao Xue, Chaitanya Baranwal, Yongbin Li 외

Transformer achieves promising results on various tasks. However, self-attention suffers from quadratic memory requirements with respect to the sequence length. Existing work focuses on reducing time and space complexity…

GPU