paper-with-me

홈 › Papers

Feedback Forensics: A Toolkit to Measure AI Personality

2025-09-30 · Arduin Findeis, Timo Kaufmann, Eyke Hüllermeier, Robert Mullins arxiv

Some traits making a "good" AI model are hard to describe upfront. For example, should responses be more polite or more casual? Such traits are sometimes summarized as model character or personality. Without a clear objective, conventional benchmarks based on automatic validation struggle to measure such traits. Evaluation methods using human feedback such as Chatbot Arena have emerged as a popular alternative. These methods infer "better" personality and other desirable traits implicitly by ranking multiple model responses relative to each other. Recent issues with model releases highlight limitations of these existing opaque evaluation approaches: a major model was rolled back over sycophantic personality issues, models were observed overfitting to such feedback-based leaderboards. Despite these known issues, limited public tooling exists to explicitly evaluate model personality. We introduce Feedback Forensics: an open-source toolkit to track AI personality changes, both those encouraged by human (or AI) feedback, and those exhibited across AI models trained and evaluated on such feedback. Leveraging AI annotators, our toolkit enables investigating personality via Python API and browser app. We demonstrate the toolkit's usefulness in two steps: (A) first we analyse the personality traits encouraged in popular human feedback datasets including Chatbot Arena, MultiPref and PRISM; and (B) then use our toolkit to analyse how much popular models exhibit such traits. We release (1) our Feedback Forensics toolkit alongside (2) a web app tracking AI personality in popular models and feedback datasets as well as (3) the underlying annotation data at https://github.com/rdnfn/feedback-forensics.

📄 PDF Abstract BibTeX arXiv:2509.26305

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Which Feedback Works for Whom? Differential Effects of LLM-Generated Feedback Elements Across Learner Profiles

2026-02-12 · Momoka Furuhashi, Kouta Nakayama, Noboru Kawai, Takashi Kodama 외 arxiv

Large language models (LLMs) show promise for automatically generating feedback in education settings. However, it remains unclear how specific feedback elements, such as tone and information coverage, contribute to lear…

A Multi-modal Personality Prediction System

2020-12-01 · ICON 2020 12 · Chanchal Suman, Aditya Gupta, Sriparna Saha, Pushpak Bhattacharyya

Automatic prediction of personality traits has many real-life applications, e.g., in forensics, recommender systems, personalized services etc.. In this work, we have proposed a solution framework for solving the problem…

Personality Trait Recognition by FacePredictionRecommendation Systems

Personality Trait Identification Using the Russian Feature Extraction Toolkit

2021-09-01 · RANLP 2021 9 · James R. Hull, Valerie Novak, C. Anton Rytting, Paul Rodrigues 외

Feature engineering is an important step in classical NLP pipelines, but machine learning engineers may not be aware of the signals to look for when processing foreign language text. The Russian Feature Extraction Toolki…

Feature EngineeringSentence

Enabling the Analysis of Personality Aspects in Recommender Systems

2020-01-07 · Shahpar Yakhchi, Amin Beheshti, Seyed Mohssen Ghafari, Mehmet Orgun

Existing Recommender Systems mainly focus on exploiting users' feedback, e.g., ratings, and reviews on common items to detect similar users. Thus, they might fail when there are no common items of interest among users. W…

Recommendation Systems

Self-Assessment Tests are Unreliable Measures of LLM Personality

2023-09-15 · Akshat Gupta, Xiaoyang Song, Gopala Anumanchipalli

As large language models (LLM) evolve in their capabilities, various recent studies have tried to quantify their behavior using psychological tools created to study human behavior. One such example is the measurement of …

Multiple-choice