paper-with-me

홈 › Papers

Test Suites Task: Evaluation of Gender Fairness in MT with MuST-SHE and INES

2023-10-30 · Beatrice Savoldi, Marco Gaido, Matteo Negri, Luisa Bentivogli

As part of the WMT-2023 "Test suites" shared task, in this paper we summarize the results of two test suites evaluations: MuST-SHE-WMT23 and INES. By focusing on the en-de and de-en language pairs, we rely on these newly created test suites to investigate systems' ability to translate feminine and masculine gender and produce gender-inclusive translations. Furthermore we discuss metrics associated with our test suites and validate them by means of human evaluations. Our results indicate that systems achieve reasonable and comparable performance in correctly translating both feminine and masculine gender forms for naturalistic gender phenomena. Instead, the generation of inclusive language forms in translation emerges as a challenging task for all the evaluated MT models, indicating room for future improvements and research on the topic.

📄 PDF Abstract BibTeX arXiv:2310.19345

Code (1)

hlt-mt/fbk-fairseq 공식 구현 pytorch

Tasks

de-enFairness

Similar Papers 제목 키워드 기반

FAIRE: Assessing Racial and Gender Bias in AI-Driven Resume Evaluations

2025-04-02 · Athena Wen, Tanush Patil, Ansh Saxena, Yicheng Fu 외

In an era where AI-driven hiring is transforming recruitment practices, concerns about fairness and bias have become increasingly important. To explore these issues, we introduce a benchmark, FAIRE (Fairness Assessment I…

Fairness

Underneath the Numbers: Quantitative and Qualitative Gender Fairness in LLMs for Depression Prediction

2024-06-12 · Micol Spitale, Jiaee Cheong, Hatice Gunes

Recent studies show bias in many machine learning models for depression detection, but bias in LLMs for this task remains unexplored. This work presents the first attempt to investigate the degree of gender bias present …

Depression DetectionFairness

Automated Evaluation of Gender Bias Across 13 Large Multimodal Models

2025-09-08 · Juan Manuel Contreras arxiv

Large multimodal models (LMMs) have revolutionized text-to-image generation, but they risk perpetuating the harmful social biases in their training data. Prior work has identified gender bias in these models, but methodo…

Text-to-Image Generation

FairQE: Multi-Agent Framework for Mitigating Gender Bias in Translation Quality Estimation

2026-04-23 · Jinhee Jang, Juhwan Choi, Dongjin Lee, Seunguk Yu 외 arxiv

Quality Estimation (QE) aims to assess machine translation quality without reference translations, but recent studies have shown that existing QE models exhibit systematic gender bias. In particular, they tend to favor m…

Machine Translation

PRIDE -- Parameter-Efficient Reduction of Identity Discrimination for Equality in LLMs

2025-07-18 · Maluna Menke, Thilo Hagendorff arxiv

Large Language Models (LLMs) frequently reproduce the gender- and sexual-identity prejudices embedded in their training corpora, leading to outputs that marginalize LGBTQIA+ users. Hence, reducing such biases is of great…

parameter-efficient fine-tuning