paper-with-me

홈 › Papers

Assessing the quality and coherence of word embeddings after SCM-based intersectional bias mitigation

2026-01-07 · Eren Kocadag, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag arxiv

Static word embeddings often absorb social biases from the text they learn from, and those biases can quietly shape downstream systems. Prior work that uses the Stereotype Content Model (SCM) has focused mostly on single-group bias along warmth and competence. We broaden that lens to intersectional bias by building compound representations for pairs of social identities through summation or concatenation, and by applying three debiasing strategies: Subtraction, Linear Projection, and Partial Projection. We study three widely used embedding families (Word2Vec, GloVe, and ConceptNet Numberbatch) and assess them with two complementary views of utility: whether local neighborhoods remain coherent and whether analogy behavior is preserved. Across models, SCM-based mitigation carries over well to the intersectional case and largely keeps the overall semantic landscape intact. The main cost is a familiar trade off: methods that most tightly preserve geometry tend to be more cautious about analogy behavior, while more assertive projections can improve analogies at the expense of strict neighborhood stability. Partial Projection is reliably conservative and keeps representations steady; Linear Projection can be more assertive; Subtraction is a simple baseline that remains competitive. The choice between summation and concatenation depends on the embedding family and the application goal. Together, these findings suggest that intersectional debiasing with SCM is practical in static embeddings, and they offer guidance for selecting aggregation and debiasing settings when balancing stability against analogy performance.

📄 PDF Abstract BibTeX arXiv:2601.04393

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BAHP: Benchmark of Assessing Word Embeddings in Historical Portuguese

2021-11-01 · EMNLP (LaTeCHCLfL, CLFL, LaTeCH) 2021 11 · Zuoyu Tian, Dylan Jarrett, Juan Escalona Torres, Patricia Amaral

High quality distributional models can capture lexical and semantic relations between words. Hence, researchers design various intrinsic tasks to test whether such relations are captured. However, most of the intrinsic t…

Outlier DetectionWord Embeddings

What do you mean, BERT? Assessing BERT as a Distributional Semantics Model

2019-11-13 · Timothee Mickus, Denis Paperno, Mathieu Constant, Kees Van Deemter

Contextualized word embeddings, i.e. vector representations for words in context, are naturally seen as an extension of previous noncontextual distributional semantic models. In this work, we focus on BERT, a deep neural…

PositionSentenceWord Embeddings

Coherence models in schizophrenia

2019-06-01 · WS 2019 6 · S Just, ra, Erik Haegert, Nora Ko{\v{r}}{\'a}nov{\'a} 외

Incoherent discourse in schizophrenia has long been recognized as a dominant symptom of the mental disorder (Bleuler, 1911/1950). Recent studies have used modern sentence and word embeddings to compute coherence metrics …

SentenceWord Embeddings

Assessing Wordnets with WordNet Embeddings

2019-07-01 · GWC 2019 7 · Ruben Branco, João Rodrigues, Chakaveh Saedi, António Branco

An effective conversion method was proposed in the literature to obtain a lexical semantic space from a lexical semantic graph, thus permitting to obtain WordNet embeddings from WordNets. In this paper, we propose the ex…

Semantic SimilaritySemantic Textual SimilarityWord Embeddings

Measuring Topic Coherence through Optimal Word Buckets

2017-04-01 · EACL 2017 4 · Nitin Ramrakhiyani, Sachin Pawar, Swapnil Hingmire, Girish Palshikar

Measuring topic quality is essential for scoring the learned topics and their subsequent use in Information Retrieval and Text classification. To measure quality of Latent Dirichlet Allocation (LDA) based topics learned …

General ClassificationInformation RetrievalRetrievaltext-classification+3