paper-with-me

홈 › Papers

Translators as Invisible Teachers of AI: Copyright, Translation Memory, and the Political Economy of Linguistic Data

2026-05-24 · Masaru Yamada arxiv

This paper examines how the labour of translators has been transformed into foundational data capital for the age of artificial intelligence (AI). Translation memories (TM) and parallel corpora preserve a one-to-one correspondence between source and target text and therefore constitute extraordinarily valuable supervised training data for machine translation. The development of statistical machine translation (SMT), neural machine translation (NMT), the Transformer architecture, and multilingual large language models (LLMs) cannot be disentangled from the accumulation of such translation data. And yet, translators' renditions have been bought as deliverables under contract, segmented as technical objects, and processed as "information analysis" data under copyright law -- losing their moral, creative, and economic attribution to the translators who produced them. The paper develops two concepts to capture this process. The first is appropriation without consumption: a mode of use in which works are not read, viewed, or listened to, but only mined for statistical features -- a use that is legitimated under Article 30-4 of the Japanese Copyright Act. The second is the invisible teacherisation of translators: the process by which translators, through the construction of translation memories, post-editing, and quality assessment, have functioned as teachers of AI without recognition as such. Drawing on the data supply chain that runs from translators through language service providers (LSPs) and platforms to model developers, on a comparative reading of Japanese, European, and United States legal frameworks, on the distinction between open and proprietary AI models, and on the premium status that human-generated data has acquired in the era of model collapse, the paper asks what translators are actually afraid of, and points toward concrete directions for redistributive design.

📄 PDF Abstract BibTeX arXiv:2605.24842

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

Integration of Machine Translation and Translation Memory: Post-Editing Efforts

2021-07-01 · TRITON 2021 7 · Rocío Caro Quintana

The development of Translation Technologies, like Translation Memory and Machine Translation, has completely changed the translation industry and translator’s workflow in the last decades. Nevertheless, TM and MT have be…

Machine TranslationTranslation

Found in Translation: Reconstructing Phylogenetic Language Trees from Translations

2017-04-24 · ACL 2017 7 · Ella Rabinovich, Noam Ordan, Shuly Wintner

Translation has played an important role in trade, law, commerce, politics, and literature for thousands of years. Translators have always tried to be invisible; ideal translations should look as if they were written ori…

Translation

TermWise: A CAT-tool with Context-Sensitive Terminological Support.

2014-05-01 · LREC 2014 5 · Kris Heylen, Stephen Bond, Dirk De Hertog, Ivan Vuli{\'c} 외

Increasingly, large bilingual document collections are being made available online, especially in the legal domain. This type of Big Data is a valuable resource that specialized translators exploit to search for informat…

Translation

Kunji : A Resource Management System for Higher Productivity in Computer Aided Translation Tools

2019-12-01 · ICON 2019 12 · Priyank Gupta, Manish Shrivastava, Dipti Misra Sharma, Rashid Ahmad

Complex NLP applications, such as machine translation systems, utilize various kinds of resources namely lexical, multiword, domain dictionaries, maps and rules etc. Similarly, translators working on Computer Aided Trans…

Machine TranslationManagementNERTranslation

Translation Memory Retrieval Using Lucene

2021-09-01 · RANLP 2021 9 · Kwang-Hyok Kim, Myong-ho Cho, Chol-ho Ryang, Ju-song Im 외

Translation Memory (TM) system, a major component of computer-assisted translation (CAT), is widely used to improve human translators’ productivity by making effective use of previously translated resource. We propose a …

Information RetrievalRetrievalTranslation