paper-with-me

Papers

Measuring Human and AI Values Based on Generative Psychometrics with Large Language Models

2024-09-18 · Haoran Ye, Yuhang Xie, Yuanyi Ren, Hanjun Fang, Xin Zhang, Guojie Song

Human values and their measurement are long-standing interdisciplinary inquiry. Recent advances in AI have sparked renewed interest in this area, with large language models (LLMs) emerging as both tools and subjects of value measurement. This work introduces Generative Psychometrics for Values (GPV), an LLM-based, data-driven value measurement paradigm, theoretically grounded in text-revealed selective perceptions. The core idea is to dynamically parse unstructured texts into perceptions akin to static stimuli in traditional psychometrics, measure the value orientations they reveal, and aggregate the results. Applying GPV to human-authored blogs, we demonstrate its stability, validity, and superiority over prior psychological tools. Then, extending GPV to LLM value measurement, we advance the current art with 1) a psychometric methodology that measures LLM values based on their scalable and free-form outputs, enabling context-specific measurement; 2) a comparative analysis of measurement paradigms, indicating response biases of prior methods; and 3) an attempt to bridge LLM values and their safety, revealing the predictive power of different value systems and the impacts of various values on LLM safety. Through interdisciplinary efforts, we aim to leverage AI for next-generation psychometrics and psychometrics for value-aligned AI.

📄 PDF Abstract BibTeX arXiv:2409.12106

Code (1)

value4ai/gpv 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement

2025-05-13 · Haoran Ye, Jing Jin, Yuhang Xie, Xin Zhang 외

The rapid advancement of large language models (LLMs) has outpaced traditional evaluation methodologies. It presents novel challenges, such as measuring human-like psychological constructs, navigating beyond static and t…

BenchmarkingLanguage ModelingLanguage ModellingLarge Language Model

PATCH! {P}sychometrics-{A}ssis{T}ed Ben{CH}marking of Large Language Models against Human Populations: A Case Study of Proficiency in 8th Grade Mathematics

2024-04-02 · Qixiang Fang, Daniel L. Oberski, Dong Nguyen

Many existing benchmarks of large (multimodal) language models (LLMs) focus on measuring LLMs' academic proficiency, often with also an interest in comparing model performance with human test takers'. While such benchmar…

Benchmarking

Undesirable Biases in NLP: Addressing Challenges of Measurement

2022-11-24 · Oskar van der Wal, Dominik Bachmann, Alina Leidinger, Leendert van Maanen 외

As Large Language Models and Natural Language Processing (NLP) technology rapidly develop and spread into daily life, it becomes crucial to anticipate how their use could harm people. One problem that has received a lot …

Quantifying AI Psychology: A Psychometrics Benchmark for Large Language Models

2024-06-25 · Yuan Li, Yue Huang, Hongyi Wang, Xiangliang Zhang 외

Large Language Models (LLMs) have demonstrated exceptional task-solving capabilities, increasingly adopting roles akin to human-like assistants. The broader integration of LLMs into society has sparked interest in whethe…

Evaluating General-Purpose AI with Psychometrics

2023-10-25 · Xiting Wang, Liming Jiang, Jose Hernandez-Orallo, David Stillwell 외

Comprehensive and accurate evaluation of general-purpose AI systems such as large language models allows for effective mitigation of their risks and deepened understanding of their capabilities. Current evaluation method…