paper-with-me

홈 › Papers

Adaptive Generation of Bias-Eliciting Questions for LLMs

2025-10-14 · Robin Staab, Jasper Dekoninck, Maximilian Baader, Martin Vechev arxiv

Large language models (LLMs) are now widely deployed in user-facing applications, reaching hundreds of millions of users worldwide. Despite their widespread adoption, growing reliance on their outputs raises significant concerns, particularly as users may be exposed to model-inherent biases that disadvantage or stereotype certain groups. However, existing bias benchmarks commonly rely on simple templated prompts or restrictive multiple-choice questions that fail to capture the complexity of real-world user interactions. In this work, we address this gap by introducing a counterfactual framework that automatically generates realistic, open-ended questions for LLM bias evaluation. Through iterative question mutation, our approach systematically explores areas where models are most likely to exhibit biased behavior. Beyond just detecting harmful biases, we also capture increasingly relevant response dimensions, such as asymmetric refusals and explicit bias acknowledgment. Building on this, we construct CAB, a diverse and human-verified benchmark for realistic and nuanced bias evaluations on current frontier LLMs. Our evaluation using CAB highlights the continued need for fairness research by showing that all examined models exhibit persistent biases across certain scenarios.

📄 PDF Abstract BibTeX arXiv:2510.12857

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BiasCause: Evaluate Socially Biased Causal Reasoning of Large Language Models

2025-04-08 · Tian Xie, Tongxin Yin, Vaishakh Keshava, Xueru Zhang 외

While large language models (LLMs) already play significant roles in society, research has shown that LLMs still generate content including social bias against certain sensitive groups. While existing benchmarks have eff…

DeCAP: Context-Adaptive Prompt Generation for Debiasing Zero-shot Question Answering in Large Language Models

2025-03-25 · Suyoung Bae, YunSeok Choi, Jee-Hyong Lee

While Large Language Models (LLMs) excel in zero-shot Question Answering (QA), they tend to expose biases in their internal knowledge when faced with socially sensitive questions, leading to a degradation in performance.…

FairnessQuestion Answering

How the Advent of Ubiquitous Large Language Models both Stymie and Turbocharge Dynamic Adversarial Question Generation

2024-01-20 · Yoo yeon Sung, Ishani Mondal, Jordan Boyd-Graber

Dynamic adversarial question generation, where humans write examples to stump a model, aims to create examples that are realistic and informative. However, the advent of large language models (LLMs) has been a double-edg…

Question GenerationQuestion-GenerationRetrieval

Eliciting Uncertainty in Chain-of-Thought to Mitigate Bias against Forecasting Harmful User Behaviors

2024-10-17 · Anthony Sicilia, Malihe Alikhani

Conversation forecasting tasks a model with predicting the outcome of an unfolding conversation. For instance, it can be applied in social media moderation to predict harmful user behaviors before they occur, allowing fo…

Eliciting Bias in Question Answering Models through Ambiguity

2021-11-01 · EMNLP (MRQA) 2021 11 · Andrew Mao, Naveen Raman, Matthew Shu, Eric Li 외

Question answering (QA) models use retriever and reader systems to answer questions. Reliance on training data by QA systems can amplify or reflect inequity through their responses. Many QA models, such as those for the …

ArticlesQuestion Answering