paper-with-me

홈 › Papers

Designing Evaluations of Machine Learning Models for Subjective Inference: The Case of Sentence Toxicity

2019-11-06 · Agathe Balayn, Alessandro Bozzon

Machine Learning (ML) is increasingly applied in real-life scenarios, raising concerns about bias in automatic decision making. We focus on bias as a notion of opinion exclusion, that stems from the direct application of traditional ML pipelines to infer subjective properties. We argue that such ML systems should be evaluated with subjectivity and bias in mind. Considering the lack of evaluation standards yet to create evaluation benchmarks, we propose an initial list of specifications to define prior to creating evaluation datasets, in order to later accurately evaluate the biases. With the example of a sentence toxicity inference system, we illustrate how the specifications support the analysis of biases related to subjectivity. We highlight difficulties in instantiating these specifications and list future work for the crowdsourcing community to help the creation of appropriate evaluation datasets.

📄 PDF Abstract BibTeX arXiv:1911.02471

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningDecision MakingSentence

Similar Papers 제목 키워드 기반

A Human-Grounded Evaluation Benchmark for Local Explanations of Machine Learning

2018-01-16 · Sina Mohseni, Jeremy E. Block, Eric D. Ragan

Research in interpretable machine learning proposes different computational and human subject approaches to evaluate model saliency explanations. These approaches measure different qualities of explanations to achieve di…

BIG-bench Machine LearningDecision MakingInterpretable Machine LearningSegmentation+1

Incorporating Eye-Tracking Signals Into Multimodal Deep Visual Models For Predicting User Aesthetic Experience In Residential Interiors

2026-01-23 · Chen-Ying Chien, Po-Chih Kuo arxiv

Understanding how people perceive and evaluate interior spaces is essential for designing environments that promote well-being. However, predicting aesthetic experiences remains difficult due to the subjective nature of …

Reproducible Subjective Evaluation

2022-03-08 · Max Morrison, Brian Tang, Gefei Tan, Bryan Pardo

Human perceptual studies are the gold standard for the evaluation of many research tasks in machine learning, linguistics, and psychology. However, these studies require significant time and cost to perform. As a result,…

The State Of TTS: A Case Study with Human Fooling Rates

2025-08-06 · Praveen Srinivasa Varadhan, Sherry Thomas, Sai Teja M. S., Suvrat Bhooshan 외 arxiv

While subjective evaluations in recent years indicate rapid progress in TTS, can current TTS systems truly pass a human deception test in a Turing-like evaluation? We introduce Human Fooling Rate (HFR), a metric that dir…

An Efficient Human Visual System Based Quality Metric for 3D Video

2018-03-13

Stereoscopic video technologies have been introduced to the consumer market in the past few years. A key factor in designing a 3D system is to understand how different visual cues and distortions affect the perceptual qu…