paper-with-me

홈 › Papers

Analyzing the Use of Character-Level Translation with Sparse and Noisy Datasets

2021-09-27 · RANLP 2013 9 · Jörg Tiedemann, Preslav Nakov

This paper provides an analysis of character-level machine translation models used in pivot-based translation when applied to sparse and noisy datasets, such as crowdsourced movie subtitles. In our experiments, we find that such character-level models cut the number of untranslated words by over 40% and are especially competitive (improvements of 2-3 BLEU points) in the case of limited training data. We explore the impact of character alignment, phrase table filtering, bitext size and the choice of pivot language on translation quality. We further compare cascaded translation models to the use of synthetic training data via multiple pivots, and we find that the latter works significantly better. Finally, we demonstrate that neither word-nor character-BLEU correlate perfectly with human judgments, due to BLEU's sensitivity to length.

📄 PDF Abstract BibTeX arXiv:2109.13723

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Noisy UGC Translation at the Character Level: Revisiting Open-Vocabulary Capabilities and Robustness of Char-Based Models

2021-10-24 · WNUT (ACL) 2021 11 · José Carlos Rosales Núñez, Guillaume Wisniewski, Djamé Seddah

This work explores the capacities of character-based Neural Machine Translation to translate noisy User-Generated Content (UGC) with a strong focus on exploring the limits of such approaches to handle productive UGC phen…

Machine TranslationTranslation

How to Learn in a Noisy World? Self-Correcting the Real-World Data Noise on Machine Translation

2024-07-02 · Yan Meng, Di wu, Christof Monz

The massive amounts of web-mined parallel data contain large amounts of noise. Semantic misalignment, as the primary source of the noise, poses a challenge for training machine translation systems. In this paper, we firs…

Machine TranslationSemantic SimilaritySemantic Textual SimilarityTranslation

Pixel-level Reconstruction and Classification for Noisy Handwritten Bangla Characters

2018-06-21 · Manohar Karki, Qun Liu, Robert DiBiano, Saikat Basu 외

Classification techniques for images of handwritten characters are susceptible to noise. Quadtrees can be an efficient representation for learning from sparse features. In this paper, we improve the effectiveness of prob…

ClassificationDocument Image ClassificationGeneral ClassificationImage Classification

Aligning Vector-spaces with Noisy Supervised Lexicon

2019-06-01 · NAACL 2019 6 · Noa Yehezkel Lubin, Jacob Goldberger, Yoav Goldberg

The problem of learning to translate between two vector spaces given a set of aligned points arises in several application areas of NLP. Current solutions assume that the lexicon which defines the alignment pairs is nois…

Translation

Aligning Vector-spaces with Noisy Supervised Lexicons

2019-03-25 · Noa Yehezkel Lubin, Jacob Goldberger, Yoav Goldberg

The problem of learning to translate between two vector spaces given a set of aligned points arises in several application areas of NLP. Current solutions assume that the lexicon which defines the alignment pairs is nois…

Translation