MiTTenS: A Dataset for Evaluating Gender Mistranslation
Translation systems, including foundation models capable of translation, can produce errors that result in gender mistranslation, and such errors can be especially harmful. To measure the extent of such potential harms when translating into and out of English, we introduce a dataset, MiTTenS, covering 26 languages from a variety of language families and scripts, including several traditionally under-represented in digital resources. The dataset is constructed with handcrafted passages that target known failure patterns, longer synthetically generated passages, and natural passages sourced from multiple domains. We demonstrate the usefulness of the dataset by evaluating both neural machine translation systems and foundation models, and show that all systems exhibit gender mistranslation and potential harm, even in high resource languages.
Code (1)
Tasks
Machine TranslationTranslationSimilar Papers 제목 키워드 기반
Mittens: An Extension of GloVe for Learning Domain-Specialized Representations
We present a simple extension of the GloVe representation learning model that begins with general-purpose representations and updates them based on data from a specialized domain. We show that the resulting representatio…
Representation LearningExploring with Sticky Mittens: Reinforcement Learning with Expert Interventions via Option Templates
Long horizon robot learning tasks with sparse rewards pose a significant challenge for current reinforcement learning algorithms. A key feature enabling humans to learn challenging control tasks is that they often receiv…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)MT-GenEval: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine Translation
As generic machine translation (MT) quality has improved, the need for targeted benchmarks that explore fine-grained aspects of quality has increased. In particular, gender accuracy in translation can have implications i…
counterfactualEthicsMachine TranslationSentence+1MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation
While multilingual large language models (LLMs) perform well on high-level tasks like translation and question answering, their ability to handle grammatical gender and morphological agreement remains underexplored. In m…
Question AnsweringEvaluating Gender Bias in the Translation of Gender-Neutral Languages into English
Machine Translation (MT) continues to improve in quality and adoption, yet the inadvertent perpetuation of gender bias remains a significant concern. Despite numerous studies into gender bias in translations from gender-…
Machine TranslationSentenceTranslation