paper-with-me

Papers

Corpus Augmentation by Sentence Segmentation for Low-Resource Neural Machine Translation

2019-05-22 · Jinyi Zhang, Tadahiro Matsumoto

Neural Machine Translation (NMT) has been proven to achieve impressive results. The NMT system translation results depend strongly on the size and quality of parallel corpora. Nevertheless, for many language pairs, no rich-resource parallel corpora exist. As described in this paper, we propose a corpus augmentation method by segmenting long sentences in a corpus using back-translation and generating pseudo-parallel sentence pairs. The experiment results of the Japanese-Chinese and Chinese-Japanese translation with Japanese-Chinese scientific paper excerpt corpus (ASPEC-JC) show that the method improves translation performance.

📄 PDF Abstract BibTeX arXiv:1905.08945

Code (0)

등록된 구현이 없습니다.

Tasks

Low Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine TranslationNMTSentenceSentence segmentationTranslation

Similar Papers 제목 키워드 기반

Parallel resources for Tunisian Arabic Dialect Translation

2020-12-01 · COLING (WANLP) 2020 12 · Saméh Kchaou, Rahma Boujelbane, Lamia Hadrich-Belguith

The difficulty of processing dialects is clearly observed in the high cost of building representative corpus, in particular for machine translation. Indeed, all machine translation systems require a huge amount and good …

Data AugmentationMachine TranslationManagementSentence+1

Data centric approach to Chinese Medical Speech Recognition

2021-10-01 · ROCLING 2021 10 · Sheng-Luen Chung, Yi-Shiuan Li, Hsien-Wei Ting

Concerning the development of Chinese medical speech recognition technology, this study re-addresses earlier encountered issues in accordance with the process of Machine Learning Engineering for Production (MLOps) from a…

Data Augmentationspeech-recognitionSpeech Recognition

Not Low-Resource Anymore: Aligner Ensembling, Batch Filtering, and New Datasets for Bengali-English Machine Translation

2020-09-20 · EMNLP 2020 11 · Tahmid Hasan, Abhik Bhattacharjee, Kazi Samin, Masum Hasan 외

Despite being the seventh most widely spoken language in the world, Bengali has received much less attention in machine translation literature due to being low in resources. Most publicly available parallel corpora for B…

Machine TranslationSentenceSentence segmentationTranslation

Language Resource Addition: Dictionary or Corpus?

2014-05-01 · LREC 2014 5 · Shinsuke Mori, Graham Neubig

In this paper, we investigate the relative effect of two strategies of language resource additions to the word segmentation problem and part-of-speech tagging problem in Japanese. The first strategy is adding entries to …

Active LearningDomain AdaptationMorphological AnalysisPart-Of-Speech Tagging+2

Data Augmentation for Machine Translation via Dependency Subtree Swapping

2023-07-13 · Attila Nagy, Dorina Petra Lakatos, Botond Barta, Patrick Nanys 외

We present a generic framework for data augmentation via dependency subtree swapping that is applicable to machine translation. We extract corresponding subtrees from the dependency parse trees of the source and target s…

Data AugmentationMachine TranslationTranslation