paper-with-me

홈 › Papers

Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems

2025-06-04 · Emma Harvey, Emily Sheng, Su Lin Blodgett, Alexandra Chouldechova, Jean Garcia-Gathright, Alexandra Olteanu, Hanna Wallach

The NLP research community has made publicly available numerous instruments for measuring representational harms caused by large language model (LLM)-based systems. These instruments have taken the form of datasets, metrics, tools, and more. In this paper, we examine the extent to which such instruments meet the needs of practitioners tasked with evaluating LLM-based systems. Via semi-structured interviews with 12 such practitioners, we find that practitioners are often unable to use publicly available instruments for measuring representational harms. We identify two types of challenges. In some cases, instruments are not useful because they do not meaningfully measure what practitioners seek to measure or are otherwise misaligned with practitioner needs. In other cases, instruments - even useful instruments - are not used by practitioners due to practical and institutional barriers impeding their uptake. Drawing on measurement theory and pragmatic measurement, we provide recommendations for addressing these challenges to better meet practitioner needs.

📄 PDF Abstract BibTeX arXiv:2506.04482

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

Computational Challenges for Polysynthetic Languages

2018-08-01 · COLING 2018 8 · Judith L. Klavans

Given advances in computational linguistic analysis of complex languages using Machine Learning as well as standard Finite State Transducers, coupled with recent efforts in language revitalization, the time was right to …

BIG-bench Machine Learning

IPQA: A Benchmark for Core Intent Identification in Personalized Question Answering

2025-10-27 · Jieyong Kim, Maryam Amirizaniani, Soojin Yoon, Dongha Lee arxiv

Intent identification serves as the foundation for generating appropriate responses in personalized question answering (PQA). However, existing benchmarks evaluate only response quality or retrieval performance without d…

Question AnsweringAnswer Selection

Improving fairness in machine learning systems: What do industry practitioners need?

2018-12-13 · Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé III, Miro Dudík 외

The potential for machine learning (ML) systems to amplify social inequities and unfairness is receiving increasing popular and academic attention. A surge of recent work has focused on the development of algorithmic too…

BIG-bench Machine LearningFairness

Understanding Machine Learning Practitioners' Data Documentation Perceptions, Needs, Challenges, and Desiderata

2022-06-06 · Amy K. Heger, Liz B. Marquis, Mihaela Vorvoreanu, Hanna Wallach 외

Data is central to the development and evaluation of machine learning (ML) models. However, the use of problematic or inappropriate datasets can result in harms when the resulting models are deployed. To encourage respon…

BIG-bench Machine Learning

PREME: Preference-based Meeting Exploration through an Interactive Questionnaire

2022-05-05 · Negar Arabzadeh, Ali Ahmadvand, Julia Kiseleva, Yang Liu 외

The recent increase in the volume of online meetings necessitates automated tools for managing and organizing the material, especially when an attendee has missed the discussion and needs assistance in quickly exploring …