Agreement-based Learning of Parallel Lexicons and Phrases from Non-Parallel Corpora
We introduce an agreement-based approach to learning parallel lexicons and phrases from non-parallel corpora. The basic idea is to encourage two asymmetric latent-variable translation models (i.e., source-to-target and target-to-source) to agree on identifying latent phrase and word alignments. The agreement is defined at both word and phrase levels. We develop a Viterbi EM algorithm for jointly training the two unidirectional models efficiently. Experiments on the Chinese-English dataset show that agreement-based learning significantly improves both alignment and translation performance.
Code (0)
등록된 구현이 없습니다.
Tasks
TranslationSimilar Papers 제목 키워드 기반
Multimodal Comparable Corpora as Resources for Extracting Parallel Data: Parallel Phrases Extraction
Introducing corpora Hlava Cor and Hlava AD: Human Label Variation in Coreference and Discourse Relations
As previous research on annotator disagreement in discourse phenomena has shown, understanding text coherence varies considerably from one individual to another. To explore this phenomenon, we created two corpora with mu…
Coreference ResolutionValidation of sub-sentential paraphrases acquired from parallel monolingual corpora
word2word: A Collection of Bilingual Lexicons for 3,564 Language Pairs
We present word2word, a publicly available dataset and an open-source Python package for cross-lingual word translations extracted from sentence-level parallel corpora. Our dataset provides top-k word translations in 3,5…
SentenceTranslationParaDetox: Detoxification with Parallel Data
We present a novel pipeline for the collection of parallel data for the detoxification task. We collect non-toxic paraphrases for over 10,000 English toxic sentences. We also show that this pipeline can be used to distil…
Sentence