paper-with-me

홈 › Papers

Agreement-based Learning of Parallel Lexicons and Phrases from Non-Parallel Corpora

2016-06-15 · ACL 2016 8 · Chunyang Liu, Yang Liu, Huanbo Luan, Maosong Sun, Heng Yu

We introduce an agreement-based approach to learning parallel lexicons and phrases from non-parallel corpora. The basic idea is to encourage two asymmetric latent-variable translation models (i.e., source-to-target and target-to-source) to agree on identifying latent phrase and word alignments. The agreement is defined at both word and phrase levels. We develop a Viterbi EM algorithm for jointly training the two unidirectional models efficiently. Experiments on the Chinese-English dataset show that agreement-based learning significantly improves both alignment and translation performance.

📄 PDF Abstract BibTeX arXiv:1606.04597

Code (0)

등록된 구현이 없습니다.

Tasks

Translation

Similar Papers 제목 키워드 기반

Multimodal Comparable Corpora as Resources for Extracting Parallel Data: Parallel Phrases Extraction

2013-10-01 · IJCNLP 2013 10 · Haithem Afli, Lo{\"\i}c Barrault, Holger Schwenk
Information RetrievalLanguage ModellingMachine TranslationSpeech Recognition

Introducing corpora Hlava Cor and Hlava AD: Human Label Variation in Coreference and Discourse Relations

2026-06-24 · Anna Nedoluzhko, Šárka Zikánová, Jiří Mírovský, Milan Straka 외 arxiv

As previous research on annotator disagreement in discourse phenomena has shown, understanding text coherence varies considerably from one individual to another. To explore this phenomenon, we created two corpora with mu…

Coreference Resolution

Validation of sub-sentential paraphrases acquired from parallel monolingual corpora

2012-04-01 · EACL 2012 4 · Houda Bouamor, Aur{\'e}lien Max, Anne Vilnat

word2word: A Collection of Bilingual Lexicons for 3,564 Language Pairs

2019-11-27 · LREC 2020 5 · Yo Joong Choe, Kyubyong Park, Dongwoo Kim

We present word2word, a publicly available dataset and an open-source Python package for cross-lingual word translations extracted from sentence-level parallel corpora. Our dataset provides top-k word translations in 3,5…

SentenceTranslation

ParaDetox: Detoxification with Parallel Data

2022-05-01 · ACL 2022 5 · Varvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy 외

We present a novel pipeline for the collection of parallel data for the detoxification task. We collect non-toxic paraphrases for over 10,000 English toxic sentences. We also show that this pipeline can be used to distil…

Sentence