paper-with-me

홈 › Papers

HumBEL: A Human-in-the-Loop Approach for Evaluating Demographic Factors of Language Models in Human-Machine Conversations

2023-05-23 · Anthony Sicilia, Jennifer C. Gates, Malihe Alikhani

While demographic factors like age and gender change the way people talk, and in particular, the way people talk to machines, there is little investigation into how large pre-trained language models (LMs) can adapt to these changes. To remedy this gap, we consider how demographic factors in LM language skills can be measured to determine compatibility with a target demographic. We suggest clinical techniques from Speech Language Pathology, which has norms for acquisition of language skills in humans. We conduct evaluation with a domain expert (i.e., a clinically licensed speech language pathologist), and also propose automated techniques to complement clinical evaluation at scale. Empirically, we focus on age, finding LM capability varies widely depending on task: GPT-3.5 mimics the ability of humans ranging from age 6-15 at tasks requiring inference, and simultaneously, outperforms a typical 21 year old at memorization. GPT-3.5 also has trouble with social language use, exhibiting less than 50% of the tested pragmatic skills. Findings affirm the importance of considering demographic alignment and conversational goals when using LMs as public-facing tools. Code, data, and a package will be available.

📄 PDF Abstract BibTeX arXiv:2305.14195

Code (1)

anthonysicilia/humbel 공식 구현

Tasks

Memorization

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
{Dispute@FaQ-s}How to file a dispute with Expedia? How to file a dispute with Expedia? To file a complaint against Expedia, first try contacting their customer service directly. You can reach them by phone at…
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Attention 설명 없음
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…

Similar Papers 제목 키워드 기반

Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech Recognition

2020-02-24 · LREC 2020 5 · Xiaolei Huang, Linzi Xing, Franck Dernoncourt, Michael J. Paul

Existing research on fairness evaluation of document classification models mainly uses synthetic monolingual data without ground truth for author demographic attributes. In this work, we assemble and publish a multilingu…

Document ClassificationFairnessHate Speech Detectionspeech-recognition+1

Identifying Factors to Help Improve Existing Decomposition-Based PMI Estimation Methods

2024-08-31 · Anna-Maria Nau, Phillip Ditto, Dawnie Wolfe Steadman, Audris Mockus

Accurately assessing the postmortem interval (PMI) is an important task in forensic science. Some of the existing techniques use regression models that use a decomposition score to predict the PMI or accumulated degree d…

Modeling Human Perspectives with Socio-Demographic Representations

2026-04-20 · Leixin Zhang, Cagri Coltekin arxiv

Humans often hold different perspectives on the same issues. In many NLP tasks, annotation disagreement can reflect valid subjective perspectives. Modeling annotator perspectives and understanding their relationship with…

Contrastive Learning

Assessing the Reliability of LLMs Annotations in the Context of Demographic Bias and Model Explanation

2025-07-17 · Hadi Mohammadi, Tina Shahedi, Pablo Mosteiro, Massimo Poesio 외 arxiv

Understanding the sources of variability in annotations is crucial for developing fair NLP systems, especially for tasks like sexism detection where demographic bias is a concern. This study investigates the extent to wh…

Talent or Luck? Evaluating Attribution Bias in Large Language Models

2025-05-28 · Chahat Raj, Mahika Banerjee, Aylin Caliskan, Antonios Anastasopoulos 외

When a student fails an exam, do we tend to blame their effort or the test's difficulty? Attribution, defined as how reasons are assigned to event outcomes, shapes perceptions, reinforces stereotypes, and influences deci…

Fairness