A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationSemantic Textual SimilaritySentenceSimilar Papers 제목 키워드 기반
Deep Neural Networks at the Service of Multilingual Parallel Sentence Extraction
Wikipedia provides an invaluable source of parallel multilingual data, which are in high demand for various sorts of linguistic inquiry, including both theoretical and practical studies. We introduce a novel end-to-end n…
Information RetrievalMachine TranslationSentenceA Deep Neural Network Approach To Parallel Sentence Extraction
Parallel sentence extraction is a task addressing the data sparsity problem found in multilingual natural language processing applications. We propose an end-to-end deep neural network approach to detect translational eq…
Feature EngineeringMachine TranslationSentenceTranslation+1Bilingual Text Extraction as Reading Comprehension
In this paper, we propose a method to extract bilingual texts automatically from noisy parallel corpora by framing the problem as a token-level span prediction, such as SQuAD-style Reading Comprehension. To extract a spa…
ArticlesReading ComprehensionSentenceTranslationExtracting Parallel Sentences with Bidirectional Recurrent Neural Networks to Improve Machine Translation
Parallel sentence extraction is a task addressing the data sparsity problem found in multilingual natural language processing applications. We propose a bidirectional recurrent neural network based approach to extract pa…
ArticlesFeature EngineeringMachine TranslationSentence+1Separating Grains from the Chaff: Using Data Filtering to Improve Multilingual Translation for Low-Resourced African Languages
We participated in the WMT 2022 Large-Scale Machine Translation Evaluation for the African Languages Shared Task. This work describes our approach, which is based on filtering the given noisy data using a sentence-pair c…
Language ModelingLanguage ModellingMachine TranslationSentence+1