paper-with-me

홈 › Papers

Uncovering Implicit Bias in Large Language Models with Concept Learning Dataset

2025-09-21 · Leroy Z. Wang arxiv

We introduce a dataset of concept learning tasks that helps uncover implicit biases in large language models. Using in-context concept learning experiments, we found that language models may have a bias toward upward monotonicity in quantifiers; such bias is less apparent when the model is tested by direct prompting without concept learning components. This demonstrates that in-context concept learning can be an effective way to discover hidden biases in language models.

📄 PDF Abstract BibTeX arXiv:2510.01219

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Uncovering Implicit Gender Bias in Narratives through Commonsense Inference

2021-09-14 · Findings (EMNLP) 2021 11 · Tenghao Huang, Faeze Brahman, Vered Shwartz, Snigdha Chaturvedi

Pre-trained language models learn socially harmful biases from their training corpora, and may repeat these biases when used for generation. We study gender biases associated with the protagonist in model-generated stori…

Actions Speak Louder than Words: Agent Decisions Reveal Implicit Biases in Language Models

2025-01-29 · YuXuan Li, Hirokazu Shirado, Sauvik Das

While advances in fairness and alignment have helped mitigate overt biases exhibited by large language models (LLMs) when explicitly prompted, we hypothesize that these models may still exhibit implicit biases when simul…

Decision MakingFairness

Can We Derive Explicit and Implicit Bias from Corpus?

2019-05-31 · Bo Wang, Baixiang Xue, Anthony G. Greenwald

Language is a popular resource to mine speakers' attitude bias, supposing that speakers' statements represent their bias on concepts. However, psychology studies show that people's explicit bias in statements can be diff…

Language Models Surface the Unwritten Code of Science and Society

2025-05-25 · Honglin Bao, Siyang Wu, Jiwoong Choi, Yingrong Mao 외

This paper calls on the research community not only to investigate how human biases are inherited by large language models (LLMs) but also to explore how these biases in LLMs can be leveraged to make society's "unwritten…

Diagnostic

Measuring Implicit Bias in Explicitly Unbiased Large Language Models

2024-02-06 · Xuechunzi Bai, Angelina Wang, Ilia Sucholutsky, Thomas L. Griffiths

Large language models (LLMs) can pass explicit social bias tests but still harbor implicit biases, similar to humans who endorse egalitarian beliefs yet exhibit subtle biases. Measuring such implicit biases can be a chal…

Decision MakingDiagnosticLanguage Modelling