paper-with-me

Papers

Masculine Defaults via Gendered Discourse in Podcasts and Large Language Models

2025-04-15 · Maria Teleki, Xiangjue Dong, Haoran Liu, James Caverlee

Masculine defaults are widely recognized as a significant type of gender bias, but they are often unseen as they are under-researched. Masculine defaults involve three key parts: (i) the cultural context, (ii) the masculine characteristics or behaviors, and (iii) the reward for, or simply acceptance of, those masculine characteristics or behaviors. In this work, we study discourse-based masculine defaults, and propose a twofold framework for (i) the large-scale discovery and analysis of gendered discourse words in spoken content via our Gendered Discourse Correlation Framework (GDCF); and (ii) the measurement of the gender bias associated with these gendered discourse words in LLMs via our Discourse Word-Embedding Association Test (D-WEAT). We focus our study on podcasts, a popular and growing form of social media, analyzing 15,117 podcast episodes. We analyze correlations between gender and discourse words -- discovered via LDA and BERTopic -- to automatically form gendered discourse word lists. We then study the prevalence of these gendered discourse words in domain-specific contexts, and find that gendered discourse-based masculine defaults exist in the domains of business, technology/politics, and video games. Next, we study the representation of these gendered discourse words from a state-of-the-art LLM embedding model from OpenAI, and find that the masculine discourse words have a more stable and robust representation than the feminine discourse words, which may result in better system performance on downstream tasks for men. Hence, men are rewarded for their discourse patterns with better system performance by one of the state-of-the-art language models -- and this embedding disparity is a representational harm and a masculine default.

📄 PDF Abstract BibTeX arXiv:2504.11431

Code (1)

mariateleki/masculine-defaults 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

LDA Linear discriminant analysis (LDA), normal discriminant analysis (NDA), or discriminant function analysis is a generalization of Fisher's linear discriminant, a method used in…
Focus 설명 없음

Similar Papers 제목 키워드 기반

Voice, Bias, and Coreference: An Interpretability Study of Gender in Speech Translation

2025-11-26 · Lina Conti, Dennis Fucci, Marco Gaido, Matteo Negri 외 arxiv

Unlike text, speech conveys information about the speaker, such as gender, through acoustic cues like pitch. This gives rise to modality-specific bias concerns. For example, in speech translation (ST), when translating f…

Transcending the "Male Code": Implicit Masculine Biases in NLP Contexts

2023-04-22 · Katie Seaborn, Shruti Chandra, Thibault Fabre

Critical scholarship has elevated the problem of gender bias in data sets used to train virtual assistants (VAs). Most work has focused on explicit biases in language, especially against women, girls, femme-identifying p…

Word Embeddings

Multilingual Holistic Bias: Extending Descriptors and Patterns to Unveil Demographic Biases in Languages at Scale

2023-05-22 · Marta R. Costa-jussà, Pierre Andrews, Eric Smith, Prangthip Hansanti 외

We introduce a multilingual extension of the HOLISTICBIAS dataset, the largest English template-based taxonomy of textual people references: MULTILINGUALHOLISTICBIAS. This extension consists of 20,459 sentences in 50 lan…

Joint Multilingual Sentence RepresentationsSentence

Gendered Language in Resumes

2021-10-16 · ACL ARR October 2021 10 · Anonymous

Despite growing concerns around gender bias in NLP models used in algorithmic hiring, there is little empirical work studying the extent and nature of gendered language in resumes. Using a corpus of 709k resumes from IT …

Gender Bias in MT for a Genderless Language: New Benchmarks for Basque

2026-03-09 · Amaia Murillo, Olatz-Perez-de-Viñaspre, Naiara Perez arxiv

Large language models (LLMs) and machine translation (MT) systems are increasingly used in our daily lives, but their outputs can reproduce gender bias present in the training data. Most resources for evaluating such bia…

Machine Translation