Evaluation of Summarization Systems across Gender, Age, and Race
Summarization systems are ultimately evaluated by human annotators and raters. Usually, annotators and raters do not reflect the demographics of end users, but are recruited through student populations or crowdsourcing platforms with skewed demographics. For two different evaluation scenarios -- evaluation against gold summaries and system output ratings -- we show that summary evaluation is sensitive to protected attributes. This can severely bias system development and evaluation, leading us to build models that cater for some groups rather than others.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Who Does the Giant Number Pile Like Best: Analyzing Fairness in Hiring Contexts
Large language models (LLMs) are increasingly being deployed in high-stakes applications like hiring, yet their potential for unfair decision-making and outcomes remains understudied, particularly in generative settings.…
Decision MakingFairnessRetrievalSensitivityExamining Gender and Race Bias in Two Hundred Sentiment Analysis Systems
Automatic machine learning systems can inadvertently accentuate and perpetuate inappropriate human biases. Past work on examining inappropriate biases has largely focused on just individual systems. Further, there is no …
Sentiment AnalysisVocal Bursts Valence PredictionDeep Generative Views to Mitigate Gender Classification Bias Across Gender-Race Groups
Published studies have suggested the bias of automated face-based gender classification algorithms across gender-race groups. Specifically, unequal accuracy rates were obtained for women and dark-skinned people. To mitig…
ClassificationFacial Attribute ClassificationFairnessGender ClassificationTowards Reducing Bias in Gender Classification
Societal bias towards certain communities is a big problem that affects a lot of machine learning systems. This work aims at addressing the racial bias present in many modern gender recognition systems. We learn race inv…
BIG-bench Machine LearningClassificationGender ClassificationGeneral ClassificationFairFace: Face Attribute Dataset for Balanced Race, Gender, and Age
Existing public face datasets are strongly biased toward Caucasian faces, and other races (e.g., Latino) are significantly underrepresented. This can lead to inconsistent model accuracy, limit the applicability of face a…
AttributeFacial Attribute Classification