paper-with-me

홈 › Papers

Towards Tailored Recovery of Lexical Diversity in Literary Machine Translation

2024-08-30 · Esther Ploeger, Huiyuan Lai, Rik van Noord, Antonio Toral

Machine translations are found to be lexically poorer than human translations. The loss of lexical diversity through MT poses an issue in the automatic translation of literature, where it matters not only what is written, but also how it is written. Current methods for increasing lexical diversity in MT are rigid. Yet, as we demonstrate, the degree of lexical diversity can vary considerably across different novels. Thus, rather than aiming for the rigid increase of lexical diversity, we reframe the task as recovering what is lost in the machine translation process. We propose a novel approach that consists of reranking translation candidates with a classifier that distinguishes between original and translated text. We evaluate our approach on 31 English-to-Dutch book translations, and find that, for certain books, our approach retrieves lexical diversity scores that are close to human translation.

📄 PDF Abstract BibTeX arXiv:2408.17308

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityMachine TranslationRerankingTranslation

Similar Papers 제목 키워드 기반

Lexical and Statistical Analysis of Bangla Newspaper and Literature: A Corpus-Driven Study on Diversity, Readability, and NLP Adaptation

2025-12-15 · Pramit Bhattacharyya, Arnab Bhattacharya arxiv

In this paper, we present a comprehensive corpus-driven analysis of Bangla literary and newspaper texts to investigate their lexical diversity, structural complexity and readability. We undertook Vacaspati and IndicCorp,…

Multi-perspective Alignment for Increasing Naturalness in Neural Machine Translation

2024-12-11 · Huiyuan Lai, Esther Ploeger, Rik van Noord, Antonio Toral

Neural machine translation (NMT) systems amplify lexical biases present in their training data, leading to artificially impoverished language in output translations. These language-level characteristics render automatic …

DiversityMachine TranslationNMTTranslation

Comparative Computational Analysis of Global Structure in Canonical, Non-Canonical and Non-Literary Texts

2020-08-25 · Mahdi Mohseni, Volker Gast, Christoph Redies

This study investigates global properties of literary and non-literary texts. Within the literary texts, a distinction is made between canonical and non-canonical works. The central hypothesis of the study is that the th…

POSSentenceTime SeriesTime Series Analysis

Deep Learning for Asynchronous Massive Access with Data Frame Length Diversity

2023-05-12 · Yanna Bai, Wei Chen, Bo Ai, Petar Popovski

Grant-free non-orthogonal multiple access has been regarded as a viable approach to accommodate access for a massive number of machine-type devices with small data packets. The sporadic activation of the devices creates …

Action DetectionActivity Detectioncompressed sensingDeep Learning+1

A Data-Oriented Model of Literary Language

2017-01-12 · EACL 2017 4 · Andreas van Cranenburgh, Rens Bod

We consider the task of predicting how literary a text is, with a gold standard from human ratings. Aside from a standard bigram baseline, we apply rich syntactic tree fragments, mined from the training set, and a series…

model