paper-with-me

Papers

Beneath the Surface: How Large Language Models Reflect Hidden Bias

2025-02-27 · Jinhao Pan, Chahat Raj, Ziyu Yao, Ziwei Zhu

The exceptional performance of Large Language Models (LLMs) often comes with the unintended propagation of social biases embedded in their training data. While existing benchmarks evaluate overt bias through direct term associations between bias concept terms and demographic terms, LLMs have become increasingly adept at avoiding biased responses, creating an illusion of neutrality. However, biases persist in subtler, contextually hidden forms that traditional benchmarks fail to capture. We introduce the Hidden Bias Benchmark (HBB), a novel dataset designed to assess hidden bias that bias concepts are hidden within naturalistic, subtly framed contexts in real-world scenarios. We analyze six state-of-the-art LLMs, revealing that while models reduce bias in response to overt bias, they continue to reinforce biases in nuanced settings. Data, code, and results are available at https://github.com/JP-25/Hidden-Bias-Benchmark.

📄 PDF Abstract BibTeX arXiv:2502.19749

Code (1)

jp-25/hidden-bias-benchmark 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs

2026-09-04 · Seogyeong Jeong, Jaehui Hwang, Dongyoon Han, Geonmo Gu 외 arxiv

Reasoning in large language models unfolds through diverse functional operations, such as problem formulation, goal decomposition, and deduction. Although these operations are explicitly distinguished in text, little is …

Connotation Lexicon: A Dash of Sentiment Beneath the Surface Meaning

2013-08-01 · ACL 2013 8 · Song Feng, Jun Seok Kang, Polina Kuznetsova, Yejin Choi
Sentiment Analysis

Capacitive Sensor Based 2D Subsurface Imaging Technology for Non Destructive Evaluation of Building Surfaces

2019-07-22

Understanding the underlying structure of building surfaces like walls and floors is essential when carrying out building maintenance and modification work. To facilitate such work, this paper introduces a capacitive sen…

D-STEER - Preference Alignment Techniques Learn to Behave, not to Believe -- Beneath the Surface, DPO as Steering Vector Perturbation in Activation Space

2025-12-03 · Samarth Raina, Saksham Aggarwal, Aman Chadha, Vinija Jain 외 arxiv

Direct Preference Optimization (DPO) has become a standard recipe for aligning large language models, yet it is still unclear what kind of change it actually induces inside the network. This paper argues that DPO does no…

Beneath (or beyond) the surface: Discovering voice-leading patterns with skip-grams

2020-06-27 · David R. W. Sears, Gerhard Widmer

Recurrent voice-leading patterns like the Mi-Re-Do compound cadence (MRDCC) rarely appear on the musical surface in complex polyphonic textures, so finding these patterns using computational methods remains a tremendous …