paper-with-me

홈 › Papers

HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment

2025-10-14 · Ali Mekky, Omar El Herraoui, Preslav Nakov, Yuxia Wang arxiv

Large language models (LLMs) are increasingly deployed across high-impact domains, from clinical decision support and legal analysis to hiring and education, making fairness and bias evaluation before deployment critical. However, existing evaluations lack grounding in real-world scenarios and do not account for differences in harm severity, e.g., a biased decision in surgery should not be weighed the same as a stylistic bias in text summarization. To address this gap, we introduce HALF (Harm-Aware LLM Fairness), a deployment-aligned framework that assesses model bias in realistic applications and weighs the outcomes by harm severity. HALF organizes nine application domains into three tiers (Severe, Moderate, Mild) using a five-stage pipeline. Our evaluation results across eight LLMs show that (1) LLMs are not consistently fair across domains, (2) model size or performance do not guarantee fairness, and (3) reasoning models perform better in medical decision support but worse in education. We conclude that HALF exposes a clear gap between previous benchmarking success and deployment readiness.

📄 PDF Abstract BibTeX arXiv:2510.12217

Code (0)

등록된 구현이 없습니다.

Tasks

Text Summarization

Similar Papers 제목 키워드 기반

Fairness Aware Reward Optimization

2026-02-08 · Ching Lam Choi, Vighnesh Subramaniam, Phillip Isola, Antonio Torralba 외 arxiv

Demographic skews in human preference data propagate systematic unfairness through reward models into aligned LLMs. We introduce Fairness Aware Reward Optimization (Faro), an in-processing framework that trains reward mo…

Can Fairness be Automated? Guidelines and Opportunities for Fairness-aware AutoML

2023-03-15 · Hilde Weerts, Florian Pfisterer, Matthias Feurer, Katharina Eggensperger 외

The field of automated machine learning (AutoML) introduces techniques that automate parts of the development of machine learning (ML) systems, accelerating the process and reducing barriers for novices. However, decisio…

AutoMLFairness

Benchmarking Bias Mitigation Toward Fairness Without Harm from Vision to LVLMs

2026-02-03 · Xuwei Tan, Ziyu Hu, Xueru Zhang arxiv

Machine learning models trained on real-world data often inherit and amplify biases against certain social groups, raising urgent concerns about their deployment at scale. While numerous bias mitigation methods have been…

mFARM: Towards Multi-Faceted Fairness Assessment based on HARMs in Clinical Decision Support

2025-09-02 · Shreyash Adappanavar, Krithi Shailya, Gokul S Krishnan, Sriraam Natarajan 외 arxiv

The deployment of Large Language Models (LLMs) in high-stakes medical settings poses a critical AI alignment challenge, as models can inherit and amplify societal biases, leading to significant disparities. Existing fair…

Bias and Fairness in Chatbots: An Overview

2023-09-16 · Jintang Xue, Yun-Cheng Wang, Chengwei Wei, Xiaofeng Liu 외

Chatbots have been studied for more than half a century. With the rapid development of natural language processing (NLP) technologies in recent years, chatbots using large language models (LLMs) have received much attent…

ChatbotFairness