paper-with-me

홈 › Papers

Towards Multi-Stakeholder Evaluation of ML Models: A Crowdsourcing Study on Metric Preferences in Job-matching System

2025-03-03 · Takuya Yokota, Yuri Nakao

While machine learning (ML) technology affects diverse stakeholders, there is no one-size-fits-all metric to evaluate the quality of outputs, including performance and fairness. Using predetermined metrics without soliciting stakeholder opinions is problematic because it leads to an unfair disregard for stakeholders in the ML pipeline. In this study, to establish practical ways to incorporate diverse stakeholder opinions into the selection of metrics for ML, we investigate participants' preferences for different metrics by using crowdsourcing. We ask 837 participants to choose a better model from two hypothetical ML models in a hypothetical job-matching system twenty times and calculate their utility values for seven metrics. To examine the participants' feedback in detail, we divide them into five clusters based on their utility values and analyze the tendencies of each cluster, including their preferences for metrics and common attributes. Based on the results, we discuss the points that should be considered when selecting appropriate metrics and evaluating ML models with multiple stakeholders.

📄 PDF Abstract BibTeX arXiv:2503.05796

Code (0)

등록된 구현이 없습니다.

Tasks

Fairness

Similar Papers 제목 키워드 기반

Crowdsourcing for Evaluating Machine Translation Quality

2014-05-01 · LREC 2014 5 · Shinsuke Goto, Donghui Lin, Toru Ishida

The recent popularity of machine translation has increased the demand for the evaluation of translations. However, the traditional evaluation approach, manual checking by a bilingual professional, is too expensive and to…

Machine TranslationSentenceTranslation

Popularity Bias in Recommendation: A Multi-stakeholder Perspective

2020-08-19 · Himan Abdollahpouri

Traditionally, especially in academic research in recommender systems, the focus has been solely on the satisfaction of the end-user. While user satisfaction has, indeed, been associated with the success of the business,…

Recommendation Systems

Power-up! What Can Generative Models Do for Human Computation Workflows?

2023-07-05 · Garrett Allen, Gaole He, Ujwal Gadiraju

We are amidst an explosion of artificial intelligence research, particularly around large language models (LLMs). These models have a range of applications across domains like medicine, finance, commonsense knowledge gra…

Knowledge Graphs

Creating Healthy Friction: Determining Stakeholder Requirements of Job Recommendation Explanations

2024-09-24 · Roan Schellingerhout, Francesco Barile, Nava Tintarev

The increased use of information retrieval in recruitment, primarily through job recommender systems (JRSs), can have a large impact on job seekers, recruiters, and companies. As a result, such systems have been determin…

FrictionInformation RetrievalRecommendation Systems

A Comparative Study on Annotation Quality of Crowdsourcing and LLM via Label Aggregation

2024-01-18 · Jiyi Li

Whether Large Language Models (LLMs) can outperform crowdsourcing on the data annotation task is attracting interest recently. Some works verified this issue with the average performance of individual crowd workers and L…