paper-with-me

홈 › Papers

BEATS: Bias Evaluation and Assessment Test Suite for Large Language Models

2025-03-31 · Alok Abhishek, Lisa Erickson, Tushar Bandopadhyay

In this research, we introduce BEATS, a novel framework for evaluating Bias, Ethics, Fairness, and Factuality in Large Language Models (LLMs). Building upon the BEATS framework, we present a bias benchmark for LLMs that measure performance across 29 distinct metrics. These metrics span a broad range of characteristics, including demographic, cognitive, and social biases, as well as measures of ethical reasoning, group fairness, and factuality related misinformation risk. These metrics enable a quantitative assessment of the extent to which LLM generated responses may perpetuate societal prejudices that reinforce or expand systemic inequities. To achieve a high score on this benchmark a LLM must show very equitable behavior in their responses, making it a rigorous standard for responsible AI evaluation. Empirical results based on data from our experiment show that, 37.65\% of outputs generated by industry leading models contained some form of bias, highlighting a substantial risk of using these models in critical decision making systems. BEATS framework and benchmark offer a scalable and statistically rigorous methodology to benchmark LLMs, diagnose factors driving biases, and develop mitigation strategies. With the BEATS framework, our goal is to help the development of more socially responsible and ethically aligned AI models.

📄 PDF Abstract BibTeX arXiv:2503.24310

Code (0)

등록된 구현이 없습니다.

Tasks

EthicsFairnessMisinformation

Similar Papers 제목 키워드 기반

Data and AI governance: Promoting equity, ethics, and fairness in large language models

2025-08-05 · Alok Abhishek, Lisa Erickson, Tushar Bandopadhyay arxiv

In this paper, we cover approaches to systematically govern, assess and quantify bias across the complete life cycle of machine learning models, from initial development and validation to ongoing production monitoring an…

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility

2026-06-11 · Xiaoyuan Liu, Jianhong Tu, Yuqi Chen, Siyuan Xie 외 arxiv

Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on fixed, LLM-centric harnesses that require heavy integration, create test-production mismatch, and limit…

Explaining Dialogue Evaluation Metrics using Adversarial Behavioral Analysis

2022-07-01 · NAACL 2022 7 · Baber Khalid, Sungjin Lee

There is an increasing trend in using neural methods for dialogue model evaluation. Lack of a framework to investigate these metrics can cause dialogue models to reflect their biases and cause unforeseen problems during …

Dialogue Evaluation

Toward Sufficient Statistical Power in Algorithmic Bias Assessment: A Test for ABROCA

2025-01-08 · Conrad Borchers

Algorithmic bias is a pressing concern in educational data mining (EDM), as it risks amplifying inequities in learning outcomes. The Area Between ROC Curves (ABROCA) metric is frequently used to measure discrepancies in …

Fairness

Promoting Generalization in Cross-Dataset Remote Photoplethysmography

2023-05-24 · Nathan Vance, Jeremy Speth, Benjamin Sporrer, Patrick Flynn

Remote Photoplethysmography (rPPG), or the remote monitoring of a subject's heart rate using a camera, has seen a shift from handcrafted techniques to deep learning models. While current solutions offer substantial perfo…