paper-with-me

홈 › Papers

Counterfactual Fairness Evaluation of LLM-Based Contact Center Agent Quality Assurance System

2026-02-16 · Kawin Mayilvaghanan, Siddhant Gupta, Ayush Kumar arxiv

Large Language Models (LLMs) are increasingly deployed in contact-center Quality Assurance (QA) to automate agent performance evaluation and coaching feedback. While LLMs offer unprecedented scalability and speed, their reliance on web-scale training data raises concerns regarding demographic and behavioral biases that may distort workforce assessment. We present a counterfactual fairness evaluation of LLM-based QA systems across 13 dimensions spanning three categories: Identity, Context, and Behavioral Style. Fairness is quantified using the Counterfactual Flip Rate (CFR), the frequency of binary judgment reversals, and the Mean Absolute Score Difference (MASD), the average shift in coaching or confidence scores across counterfactual pairs. Evaluating 18 LLMs on 3,000 real-world contact center transcripts, we find systematic disparities, with CFR ranging from 5.4% to 13.0% and consistent MASD shifts across confidence, positive, and improvement scores. Larger, more strongly aligned models show lower unfairness, though fairness does not track accuracy. Contextual priming of historical performance induces the most severe degradations (CFR up to 16.4%), while implicit linguistic identity cues remain a persistent bias source. Finally, we analyze the efficacy of fairness-aware prompting, finding that explicit instructions yield only modest improvements in evaluative consistency. Our findings underscore the need for standardized fairness auditing pipelines prior to deploying LLMs in high-stakes workforce evaluation.

📄 PDF Abstract BibTeX arXiv:2602.14970

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LLM-Based Insight Extraction for Contact Center Analytics and Cost-Efficient Deployment

2025-03-24 · Varsha Embar, Ritvik Shrivastava, Vinay Damodaran, Travis Mehlinger 외

Large Language Models have transformed the Contact Center industry, manifesting in enhanced self-service tools, streamlined administrative processes, and augmented agent productivity. This paper delineates our system tha…

AI Coach Assist: An Automated Approach for Call Recommendation in Contact Centers for Agent Coaching

2023-05-28 · Md Tahmid Rahman Laskar, Cheng Chen, Xue-Yong Fu, Mahsa Azizi 외

In recent years, the utilization of Artificial Intelligence (AI) in the contact center industry is on the rise. One area where AI can have a significant impact is in the coaching of contact center agents. By analyzing ca…

Counterfactual Fairness Evaluation of Machine Learning Models on Educational Datasets

2025-04-15 · Woojin Kim, Hyeoncheol Kim

As machine learning models are increasingly used in educational settings, from detecting at-risk students to predicting student performance, algorithmic bias and its potential impacts on students raise critical concerns …

counterfactualFairness

Towards counterfactual fairness through auxiliary variables

2024-12-06 · Bowei Tian, Ziyao Wang, Shwai He, Wanghao Ye 외

The challenge of balancing fairness and predictive accuracy in machine learning models, especially when sensitive attributes such as race, gender, or age are considered, has motivated substantial research in recent years…

counterfactualFairness

Towards Fairness Assessment of Dutch Hate Speech Detection

2025-06-14 · Julie Bauer, Rishabh Kaushal, Thales Bertaglia, Adriana Iamnitchi

Numerous studies have proposed computational methods to detect hate speech online, yet most focus on the English language and emphasize model development. In this study, we evaluate the counterfactual fairness of hate sp…

counterfactualFairnessHate Speech DetectionSentence