Announcing CzEng 2.0 Parallel Corpus with over 2 Gigawords
We present a new release of the Czech-English parallel corpus CzEng 2.0 consisting of over 2 billion words (2 "gigawords") in each language. The corpus contains document-level information and is filtered with several techniques to lower the amount of noise. In addition to the data in the previous version of CzEng, it contains new authentic and also high-quality synthetic parallel data. CzEng is freely available for research and educational purposes.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
The Joy of Parallelism with CzEng 1.0
CzEng 1.0 is an updated release of our Czech-English parallel corpus, freely available for non-commercial research or educational purposes. In this release, we approximately doubled the corpus size, reaching 15 million s…
Machine TranslationSentenceJoint search in a bilingual valency lexicon and an annotated corpus
In this paper and the associated system demo, we present an advanced search system that allows to perform a joint search over a (bilingual) valency lexicon and a correspondingly annotated linked parallel corpus. This sea…
Announcing Prague Czech-English Dependency Treebank 2.0
We introduce a substantial update of the Prague Czech-English Dependency Treebank, a parallel corpus manually annotated at the deep syntactic layer of linguistic representation. The English part consists of the Wall Stre…
Coreference ResolutionSentenceSynSemClass Linked Lexicon: Mapping Synonymy between Languages
This paper reports on an extended version of a synonym verb class lexicon, newly called SynSemClass (formerly CzEngClass). This lexicon stores cross-lingual semantically similar verb senses in synonym classes extracted f…
Synonymy in Bilingual Context: The CzEngClass Lexicon
This paper describes CzEngClass, a bilingual lexical resource being built to investigate verbal synonymy in bilingual context and to relate semantic roles common to one synonym class to verb arguments (verb valency). In …
Word Sense Disambiguation