paper-with-me

홈 › Papers

The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models

2024-04-24 · Hannah Rose Kirk, Alexander Whitefield, Paul Röttger, Andrew Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, Max Bartolo, Adina Williams, He He, Bertie Vidgen, Scott A. Hale

Human feedback is central to the alignment of Large Language Models (LLMs). However, open questions remain about methods (how), domains (where), people (who) and objectives (to what end) of feedback processes. To navigate these questions, we introduce PRISM, a dataset that maps the sociodemographics and stated preferences of 1,500 diverse participants from 75 countries, to their contextual preferences and fine-grained feedback in 8,011 live conversations with 21 LLMs. With PRISM, we contribute (i) wider geographic and demographic participation in feedback; (ii) census-representative samples for two countries (UK, US); and (iii) individualised ratings that link to detailed participant profiles, permitting personalisation and attribution of sample artefacts. We target subjective and multicultural perspectives on value-laden and controversial issues, where we expect interpersonal and cross-cultural disagreement. We use PRISM in three case studies to demonstrate the need for careful consideration of which humans provide what alignment data.

📄 PDF Abstract BibTeX arXiv:2404.16019

Code (1)

hannahkirk/prism-alignment 공식 구현

Tasks

DiversityNavigate

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Focus 설명 없음

Similar Papers 제목 키워드 기반

Adaptive Recruitment Resource Allocation to Improve Cohort Representativeness in Participatory Biomedical Datasets

2024-08-02 · Victor Borza, Andrew Estornell, Ellen Wright Clayton, Chien-Ju Ho 외

Large participatory biomedical studies, studies that recruit individuals to join a dataset, are gaining popularity and investment, especially for analysis by modern AI methods. Because they purposively recruit participan…

What are you optimizing for? Aligning Recommender Systems with Human Values

2021-07-22 · Jonathan Stray, Ivan Vendrov, Jeremy Nixon, Steven Adler 외

We describe cases where real recommender systems were modified in the service of various human values such as diversity, fairness, well-being, time well spent, and factual accuracy. From this we identify the current prac…

DiversityFairnessRecommendation Systems

What Do People Actually Want From AI? Mapping Preference Plurality

2026-06-04 · Julia Sepúlveda Coelho, Scott A. Hale arxiv

Large Language Models (LLMs) are often fine-tuned through Reinforcement Learning from Human Feedback (RLHF) to align with people's preferences and values. However, this method has known limitations: it aggregates conflic…

Reinforcement Learning

The Participatory Turn in AI Design: Theoretical Foundations and the Current State of Practice

2023-10-02 · Fernando Delgado, Stephen Yang, Michael Madaio, Qian Yang

Despite the growing consensus that stakeholders affected by AI systems should participate in their design, enormous variation and implicit disagreements exist among current approaches. For researchers and practitioners w…

Expansive Participatory AI: Supporting Dreaming within Inequitable Institutions

2022-11-22 · Michael Alan Chang, Shiran Dudy

Participatory Artificial Intelligence (PAI) has recently gained interest by researchers as means to inform the design of technology through collective's lived experience. PAI has a greater promise than that of providing …