paper-with-me

홈 › Papers

Using English as Pivot to Extract Persian-Italian Parallel Sentences from Non-Parallel Corpora

2017-01-29 · Ebrahim Ansari, M. H. Sadreddini, Mostafa Sheikhalishahi, Richard Wallace, Fatemeh Alimardani

The effectiveness of a statistical machine translation system (SMT) is very dependent upon the amount of parallel corpus used in the training phase. For low-resource language pairs there are not enough parallel corpora to build an accurate SMT. In this paper, a novel approach is presented to extract bilingual Persian-Italian parallel sentences from a non-parallel (comparable) corpus. In this study, English is used as the pivot language to compute the matching scores between source and target sentences and candidate selection phase. Additionally, a new monolingual sentence similarity metric, Normalized Google Distance (NGD) is proposed to improve the matching process. Moreover, some extensions of the baseline system are applied to improve the quality of extracted sentences measured with BLEU. Experimental results show that using the new pivot based extraction can increase the quality of bilingual corpus significantly and consequently improves the performance of the Persian-Italian SMT system.

📄 PDF Abstract BibTeX arXiv:1701.08339

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentenceSentence SimilarityTranslation

Similar Papers 제목 키워드 기반

Extracting Bilingual Persian Italian Lexicon from Comparable Corpora Using Different Types of Seed Dictionaries

2017-01-29 · Ebrahim Ansari, M. H. Sadreddini, Lucio Grandinetti, Mahsa Radinmehr 외

Bilingual dictionaries are very important in various fields of natural language processing. In recent years, research on extracting new bilingual lexicons from non-parallel (comparable) corpora have been proposed. Almost…

Creation of comparable corpora for English-Urdu, Arabic, Persian

2016-05-01 · LREC 2016 5 · Murad Abouammoh, Kashif Shah, Ahmet Aker

Statistical Machine Translation (SMT) relies on the availability of rich parallel corpora. However, in the case of under-resourced languages or some specific domains, parallel corpora are not readily available. This lead…

ArticlesMachine TranslationTranslation

Persian-Spanish Low-Resource Statistical Machine Translation Through English as Pivot Language

2017-09-01 · RANLP 2017 9 · Benyamin Ahmadnia, Javier Serrano, Gholamreza Haffari

This paper is an attempt to exclusively focus on investigating the pivot language technique in which a bridging language is utilized to increase the quality of the Persian-Spanish low-resource Statistical Machine Transla…

Machine TranslationSentenceTranslation

Extracting an English-Persian Parallel Corpus from Comparable Corpora

2017-11-02 · LREC 2018 5 · Akbar Karimi, Ebrahim Ansari, Bahram Sadeghi Bigham

Parallel data are an important part of a reliable Statistical Machine Translation (SMT) system. The more of these data are available, the better the quality of the SMT system. However, for some language pairs such as Per…

Machine TranslationTranslation

Improving Direct Persian-English Speech-to-Speech Translation with Discrete Units and Synthetic Parallel Data

2025-11-16 · Sina Rashidi, Hossein Sameti arxiv

Direct speech-to-speech translation (S2ST), in which all components are trained jointly, is an attractive alternative to cascaded systems because it offers a simpler pipeline and lower inference latency. However, direct …

Speech-to-Speech Translation