paper-with-me

Papers

Reducing the Search Space for Parallel Sentences in Comparable Corpora

2020-05-01 · LREC 2020 5 · R{\'e}mi Cardon, Natalia Grabar

This paper describes and evaluates simple techniques for reducing the research space for parallel sentences in monolingual comparable corpora. Initially, when searching for parallel sentences between two comparable documents, all the possible sentence pairs between the documents have to be considered, which introduces a great degree of imbalance between parallel pairs and non-parallel pairs. This is a problem because even with a high performing algorithm, a lot of noise will be present in the extracted results, thus introducing a need for an extensive and costly manual check phase. We work on a manually annotated subset obtained from a French comparable corpus and show how we can drastically reduce the number of sentence pairs that have to be fed to a classifier so that the results can be manually handled.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

zNLP: Identifying Parallel Sentences in Chinese-English Comparable Corpora

2017-08-01 · WS 2017 8 · Zheng Zhang, Pierre Zweigenbaum

This paper describes the zNLP system for the BUCC 2017 shared task. Our system identifies parallel sentence pairs in Chinese-English comparable corpora by translating word-by-word Chinese sentences into English, using th…

Machine TranslationSentence

A Corpus for English-Japanese Multimodal Neural Machine Translation with Comparable Sentences

2020-10-17 · Andrew Merritt, Chenhui Chu, Yuki Arase

Multimodal neural machine translation (NMT) has become an increasingly important area of research over the years because additional modalities, such as image data, can provide more context to textual data. Furthermore, t…

Image CaptioningMachine TranslationNMTSentence+1

Building Subject-aligned Comparable Corpora and Mining it for Truly Parallel Sentence Pairs

2015-09-29 · Krzysztof Wołk, Krzysztof Marasek

Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our methodology for mining such…

ArticlesMachine TranslationRetrievalSentence+1

Harvesting comparable corpora and mining them for equivalent bilingual sentences using statistical classification and analogy- based heuristics

2015-11-18 · Krzysztof Wołk, Emilia Rejmund, Krzysztof Marasek

Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our new methodologies for minin…

General ClassificationMachine TranslationRetrievalTranslation

Multi-domain machine translation enhancements by parallel data extraction from comparable corpora

2016-03-22 · Krzysztof Wołk, Emilia Rejmund, Krzysztof Marasek

Parallel texts are a relatively rare language resource, however, they constitute a very useful research material with a wide range of applications. This study presents and analyses new methodologies we developed for obta…

Machine TranslationTranslation