paper-with-me

홈 › Papers

Robustly Improving LLM Fairness in Realistic Settings via Interpretability

2025-06-12 · Adam Karvonen, Samuel Marks

Large language models (LLMs) are increasingly deployed in high-stakes hiring applications, making decisions that directly impact people's careers and livelihoods. While prior studies suggest simple anti-bias prompts can eliminate demographic biases in controlled evaluations, we find these mitigations fail when realistic contextual details are introduced. We address these failures through internal bias mitigation: by identifying and neutralizing sensitive attribute directions within model activations, we achieve robust bias reduction across all tested scenarios. Across leading commercial (GPT-4o, Claude 4 Sonnet, Gemini 2.5 Flash) and open-source models (Gemma-2 27B, Gemma-3, Mistral-24B), we find that adding realistic context such as company names, culture descriptions from public careers pages, and selective hiring constraints (e.g.,``only accept candidates in the top 10\%") induces significant racial and gender biases (up to 12\% differences in interview rates). When these biases emerge, they consistently favor Black over White candidates and female over male candidates across all tested models and scenarios. Moreover, models can infer demographics and become biased from subtle cues like college affiliations, with these biases remaining invisible even when inspecting the model's chain-of-thought reasoning. To address these limitations, our internal bias mitigation identifies race and gender-correlated directions and applies affine concept editing at inference time. Despite using directions from a simple synthetic dataset, the intervention generalizes robustly, consistently reducing bias to very low levels (typically under 1\%, always below 2.5\%) while largely maintaining model performance. Our findings suggest that practitioners deploying LLMs for hiring should adopt more realistic evaluation methodologies and consider internal mitigation strategies for equitable outcomes.

📄 PDF Abstract BibTeX arXiv:2506.10922

Code (1)

adamkarvonen/llm_bias 공식 구현 jax

Tasks

AttributeFairness

Methods 이 논문이 사용한 방법론

ADOPT Please enter a description about the method here

Similar Papers 제목 키워드 기반

Interpretable Fair Clustering

2025-11-26 · Mudi Jiang, Jiahui Zhou, Xinying Liu, Zengyou He 외 arxiv

Fair clustering has gained increasing attention in recent years, especially in applications involving socially sensitive attributes. However, existing fair clustering methods often lack interpretability, limiting their a…

Reinforcement Learning Towards Broadly and Persistently Beneficial Models

2026-06-22 · Akshay V. Jagadeesh, Rahul K. Arora, Khaled Saab, Ali Malik 외 arxiv

As AI systems are deployed across increasingly diverse and high-stakes settings, model alignment must generalize beyond the tasks and domains seen during training. This is especially important for reinforcement learning …

Reinforcement Learning

When Interpretability Is Unequally Distributed: Fairness in Hybrid Interpretable Models

2026-05-27 · Ziba Jabbar Zare, Ulrich Aïvodji, Julien Ferry, Thibaut Vidal arxiv

Hybrid interpretable models combine a transparent component with a black-box model by assigning some examples to the former and deferring the rest to the latter. While this design enables flexible tradeoffs between accur…

Federated Unlearning in the Wild: Rethinking Fairness and Data Discrepancy

2025-10-08 · ZiHeng Huang, Di Wu, Jun Bai, Jiale Zhang 외 arxiv

Machine unlearning is critical for enforcing data deletion rights like the "right to be forgotten." As a decentralized paradigm, Federated Learning (FL) also requires unlearning, but realistic implementations face two ma…

Federated Learning

Rethinking Fairness for Human-AI Collaboration

2023-10-05 · Haosen Ge, Hamsa Bastani, Osbert Bastani

Existing approaches to algorithmic fairness aim to ensure equitable outcomes if human decision-makers comply perfectly with algorithmic decisions. However, perfect compliance with the algorithm is rarely a reality or eve…

Fairness