paper-with-me

홈 › Papers

In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores

2026-04-21 · Zeyu Tang, Sang T. Truong, Deonna Owens, Shreyas Sharma, Yibo Jacky Zhang, Brando Miranda, Sanmi Koyejo arxiv

LLM fairness should be evaluated through in-situ conversational behavior rather than standardized-test Q&A benchmarks. We show that the standardized-test paradigm can be structurally unreliable: surface-level prompt construction choices, although entirely orthogonal to the fairness question being tested, account for the majority of score variance, shift fairness conclusions in both the direction and the magnitude, and result in severe discordance in model rankings. We develop MAC-Fairness, a multi-agent conversational framework that embeds controlled variation factors into multi-round dialogue for in-situ behavior evaluation, examining how models' conversational behavior shifts when identity is varied as part of natural multi-agent interaction. Repurposing standardized-test questions as conversation seeds rather than as the evaluation instrument, we evaluate position persistence (how they hold positions, from the self-perspective) and peer receptiveness (how receptive they are to peers, from the other-perspective) across 8 million conversation transcripts spanning multiple models and identity presence configurations. In-situ behavioral evaluation reveals stable, model-specific behavioral signatures that could generalize across benchmarks differing in fairness targets and evaluation methodologies, a form of evidence the standardized-test paradigm does not offer.

📄 PDF Abstract BibTeX arXiv:2605.12530

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Counterfactual Fairness Evaluation of LLM-Based Contact Center Agent Quality Assurance System

2026-02-16 · Kawin Mayilvaghanan, Siddhant Gupta, Ayush Kumar arxiv

Large Language Models (LLMs) are increasingly deployed in contact-center Quality Assurance (QA) to automate agent performance evaluation and coaching feedback. While LLMs offer unprecedented scalability and speed, their …

A Turing Test: Are AI Chatbots Behaviorally Similar to Humans?

2023-11-19 · Qiaozhu Mei, Yutong Xie, Walter Yuan, Matthew O. Jackson

We administer a Turing Test to AI Chatbots. We examine how Chatbots behave in a suite of classic behavioral games that are designed to elicit characteristics such as trust, fairness, risk-aversion, cooperation, \textit{e…

Fairness

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

2025-10-25 · Alexandra Yost, Shreyans Jain, Shivam Raval, Grant Corser 외 arxiv

Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure or superficial variation. We propose a framework to measure consistent be…

FairDiverse: A Comprehensive Toolkit for Fair and Diverse Information Retrieval Algorithms

2025-02-17 · Chen Xu, Zhirui Deng, Clara Rus, Xiaopeng Ye 외

In modern information retrieval (IR). achieving more than just accuracy is essential to sustaining a healthy ecosystem, especially when addressing fairness and diversity considerations. To meet these needs, various datas…

DiversityFairnessInformation RetrievalRetrieval

DHAuDS: A Dynamic and Heterogeneous Audio Benchmark for Test-Time Adaptation

2025-11-23 · Weichuang Shao, Iman Yi Liao, Tomas Henrique Bode Maul, Tissa Chandesa arxiv

Existing Test-time Adaptation (TTA) studies rely heavily on static and homogeneous corruption protocols, such as ImageNet-C and CIFAR-10-C/100-C, leading to inconsistent evaluation settings and potentially inflated robus…

Test-time AdaptationAudio Classification