paper-with-me

Papers

Towards Massive Multilingual Holistic Bias

2024-06-29 · Xiaoqing Ellen Tan, Prangthip Hansanti, Carleigh Wood, Bokai Yu, Christophe Ropers, Marta R. Costa-jussà

In the current landscape of automatic language generation, there is a need to understand, evaluate, and mitigate demographic biases as existing models are becoming increasingly multilingual. To address this, we present the initial eight languages from the MASSIVE MULTILINGUAL HOLISTICBIAS (MMHB) dataset and benchmark consisting of approximately 6 million sentences representing 13 demographic axes. We propose an automatic construction methodology to further scale up MMHB sentences in terms of both language coverage and size, leveraging limited human annotation. Our approach utilizes placeholders in multilingual sentence construction and employs a systematic method to independently translate sentence patterns, nouns, and descriptors. Combined with human translation, this technique carefully designs placeholders to dynamically generate multiple sentence variations and significantly reduces the human translation workload. The translation process has been meticulously conducted to avoid an English-centric perspective and include all necessary morphological variations for languages that require them, improving from the original English HOLISTICBIAS. Finally, we utilize MMHB to report results on gender bias and added toxicity in machine translation tasks. On the gender analysis, MMHB unveils: (1) a lack of gender robustness showing almost +4 chrf points in average for masculine semantic sentences compared to feminine ones and (2) a preference to overgeneralize to masculine forms by reporting more than +12 chrf points in average when evaluating with masculine compared to feminine references. MMHB triggers added toxicity up to 2.3%.

📄 PDF Abstract BibTeX arXiv:2407.00486

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentenceText GenerationTranslation

Similar Papers 제목 키워드 기반

Multilingual Holistic Bias: Extending Descriptors and Patterns to Unveil Demographic Biases in Languages at Scale

2023-05-22 · Marta R. Costa-jussà, Pierre Andrews, Eric Smith, Prangthip Hansanti 외

We introduce a multilingual extension of the HOLISTICBIAS dataset, the largest English template-based taxonomy of textual people references: MULTILINGUALHOLISTICBIAS. This extension consists of 20,459 sentences in 50 lan…

Joint Multilingual Sentence RepresentationsSentence

On Evaluating and Mitigating Gender Biases in Multilingual Settings

2023-07-04 · Aniket Vashishtha, Kabir Ahuja, Sunayana Sitaram

While understanding and removing gender biases in language models has been a long-standing problem in Natural Language Processing, prior research work has primarily been limited to English. In this work, we investigate s…

VHELM: A Holistic Evaluation of Vision Language Models

2024-10-09 · Tony Lee, Haoqin Tu, Chi Heem Wong, Wenhao Zheng 외

Current benchmarks for assessing vision-language models (VLMs) often focus on their perception or problem-solving capabilities and neglect other critical aspects such as fairness, multilinguality, or toxicity. Furthermor…

Fairness

The Balancing Act: Unmasking and Alleviating ASR Biases in Portuguese

2024-02-12 · Ajinkya Kulkarni, Anna Tokareva, Rameez Qureshi, Miguel Couceiro

In the field of spoken language understanding, systems like Whisper and Multilingual Massive Speech (MMS) have shown state-of-the-art performances. This study is dedicated to a comprehensive exploration of the Whisper an…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond

2024-08-07 · Beomseok Lee, Ioan Calapodescu, Marco Gaido, Matteo Negri 외

We present Speech-MASSIVE, a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages from different famil…

BenchmarkingLanguage Identificationslot-fillingSlot Filling+1