paper-with-me

홈 › Papers

Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora

2026-05-21 · Maciej Skorski arxiv

Moral language is subtle and culturally variable, making it difficult to translate faithfully across languages. Idiomatic expressions, slang, and cultural references introduce hard-to-avoid translation artifacts. Yet automated moral classification depends on language-specific annotated corpora that exist almost exclusively in English. We investigate whether LLM-based translation can bridge this gap, taking Polish as a test case. Using $\sim$50k morally annotated social media posts from a diverse range of topics, we apply a principled four-method validation pipeline: LaBSE cross-lingual embedding similarity, Centred Kernel Alignment (CKA), LLM-as-judge evaluation, and deep learning classifier parity tests. We show that despite shortcomings in handling slang, vulgarity, and culturally loaded expressions, direct translation preserves subtle moral cues well enough to be harvested by cross-lingual machine learning --- with a mean cosine similarity of 0.89 and classification accuracy gaps of 0.01--0.02 AUROC across foundations. These results demonstrate that machine translation is a practical and cost-effective path to moral values research in languages currently under-resourced in this domain. We demonstrate this for Polish as a representative Slavic language, with expected generalization to related languages.

📄 PDF Abstract BibTeX arXiv:2605.22660

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

Beyond English: Evaluating Automated Measurement of Moral Foundations in Non-English Discourse with a Chinese Case Study

2025-02-04 · Calvin Yixiang Cheng, Scott A Hale

This study explores computational approaches for measuring moral foundations (MFs) in non-English corpora. Since most resources are developed primarily for English, cross-linguistic applications of moral foundation theor…

Machine TranslationTransfer Learning

Does Summary Evaluation Survive Translation to Other Languages?

2021-09-16 · NAACL 2022 7 · Spencer Braun, Oleg Vasilyev, Neslihan Iskender, John Bohannon

The creation of a quality summarization dataset is an expensive, time-consuming effort, requiring the production and evaluation of summaries by both trained humans and machines. If such effort is made in one language, it…

Machine TranslationTranslation

Does Summary Evaluation Survive Translation to Other Languages?

2022-01-16 · ACL ARR January 2022 1 · Anonymous

The creation of a quality summarization dataset is an expensive, time-consuming effort, requiring the production and evaluation of summaries by both trained humans and machines. The returns to such an effort would increa…

Machine TranslationTranslation

Lost in Translation? A Comparative Study on the Cross-Lingual Transfer of Composite Harms

2026-02-08 · Vaibhav Shukla, Hardik Sharma, Adith N Reganti, Soham Wasmatkar 외 arxiv

Most safety evaluations of large language models (LLMs) remain anchored in English. Translation is often used as a shortcut to probe multilingual behavior, but it rarely captures the full picture, especially when harmful…

Cross-Lingual Transfer

Preparing an Endangered Language for the Digital Age: The Case of Judeo-Spanish

2022-05-31 · EURALI (LREC) 2022 6 · Alp Öktem, Rodolfo Zevallos, Yasmin Moslem, Güneş Öztürk 외

We develop machine translation and speech synthesis systems to complement the efforts of revitalizing Judeo-Spanish, the exiled language of Sephardic Jews, which survived for centuries, but now faces the threat of extinc…

Machine TranslationSpeech Synthesistext-to-speechText to Speech+2