paper-with-me

홈 › Papers

Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles

2025-11-08 · Fatima Jahara, Mark Dredze, Sharon Levy arxiv

While recent safety guardrails effectively suppress overtly biased outputs, subtler forms of social bias emerge during complex logical reasoning tasks that evade current evaluation benchmarks. To fill this gap, we introduce a new evaluation framework, PRIME (Puzzle Reasoning for Implicit Biases in Model Evaluation), that uses logic grid puzzles to systematically probe the influence of social stereotypes on logical reasoning and decision making in LLMs. Our use of logic puzzles enables automatic generation and verification, as well as variability in complexity and biased settings. PRIME includes stereotypical, anti-stereotypical, and neutral puzzle variants generated from a shared puzzle structure, allowing for controlled and fine-grained comparisons. We evaluate multiple model families across puzzle sizes and test the effectiveness of prompt-based mitigation strategies. Focusing our experiments on gender stereotypes, our findings highlight that models consistently reason more accurately when solutions align with stereotypical associations. This demonstrates the significance of PRIME for diagnosing and quantifying social biases perpetuated in the deductive reasoning of LLMs, where fairness is critical.

📄 PDF Abstract BibTeX arXiv:2511.06160

Code (0)

등록된 구현이 없습니다.

Tasks

Logical ReasoningDecision Making

Similar Papers 제목 키워드 기반

Position: On the Methodological Pitfalls of Evaluating Base LLMs for Reasoning

2025-11-13 · Jason Chan, Zhixue Zhao, Robert Gaizauskas arxiv

Existing work investigates the reasoning capabilities of large language models (LLMs) to uncover their limitations, human-like biases and underlying processes. Such studies include evaluations of base LLMs (pre-trained o…

Evaluating Large Language Models with NeuBAROCO: Syllogistic Reasoning Ability and Human-like Biases

2023-06-21 · Risako Ando, Takanobu Morishita, Hirohiko Abe, Koji Mineshima 외

This paper investigates whether current large language models exhibit biases in logical reasoning, similar to humans. Specifically, we focus on syllogistic reasoning, a well-studied form of inference in the cognitive sci…

Logical Reasoning

Explicit vs. Implicit: Investigating Social Bias in Large Language Models through Self-Reflection

2025-01-04 · Yachao Zhao, Bo wang, Yan Wang

Large Language Models (LLMs) have been shown to exhibit various biases and stereotypes in their generated content. While extensive research has investigated bias in LLMs, prior work has predominantly focused on explicit …

ORCHARD: A Benchmark For Measuring Systematic Generalization of Multi-Hierarchical Reasoning

2021-11-28 · Bill Tuck Weng Pung, Alvin Chan

The ability to reason with multiple hierarchical structures is an attractive and desirable property of sequential inductive biases for natural language processing. Do the state-of-the-art Transformers and LSTM architectu…

DiagnosticListOpsRelational ReasoningSystematic Generalization

ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues

2026-04-02 · Bhaskara Hanuma Vedula, Darshan Anghan, Ishita Goyal, Ponnurangam Kumaraguru 외 arxiv

Large Language Models increasingly suppress biased outputs when demographic identity is stated explicitly, yet may still exhibit implicit biases when identity is conveyed indirectly. Existing benchmarks use name based pr…