paper-with-me

홈 › Papers

Comparing Humans and Models on a Similar Scale: Towards Cognitive Gender Bias Evaluation in Coreference Resolution

2023-05-24 · Gili Lior, Gabriel Stanovsky

Spurious correlations were found to be an important factor explaining model performance in various NLP tasks (e.g., gender or racial artifacts), often considered to be ''shortcuts'' to the actual task. However, humans tend to similarly make quick (and sometimes wrong) predictions based on societal and cognitive presuppositions. In this work we address the question: can we quantify the extent to which model biases reflect human behaviour? Answering this question will help shed light on model performance and provide meaningful comparisons against humans. We approach this question through the lens of the dual-process theory for human decision-making. This theory differentiates between an automatic unconscious (and sometimes biased) ''fast system'' and a ''slow system'', which when triggered may revisit earlier automatic reactions. We make several observations from two crowdsourcing experiments of gender bias in coreference resolution, using self-paced reading to study the ''fast'' system, and question answering to study the ''slow'' system under a constrained time setting. On real-world data humans make $\sim$3\% more gender-biased decisions compared to models, while on synthetic data models are $\sim$12\% more biased.

📄 PDF Abstract BibTeX arXiv:2305.15389

Code (1)

slab-nlp/cog-gb-eval 공식 구현

Tasks

coreference-resolutionCoreference ResolutionDecision MakingQuestion Answering

Similar Papers 제목 키워드 기반

Robusto-1 Dataset: Comparing Humans and VLMs on real out-of-distribution Autonomous Driving VQA from Peru

2025-03-10 · Dunant Cusipuma, David Ortega, Victor Flores-Benites, Arturo Deza

As multimodal foundational models start being deployed experimentally in Self-Driving cars, a reasonable question we ask ourselves is how similar to humans do these systems respond in certain driving situations -- especi…

Autonomous DrivingQuestion AnsweringSelf-Driving CarsVisual Question Answering+1

The "LLM World of Words" English free association norms generated by large language models

2024-12-02 · Katherine Abramski, Riccardo Improta, Giulio Rossetti, Massimo Stella

Free associations have been extensively used in cognitive psychology and linguistics for studying how conceptual knowledge is organized. Recently, the potential of applying a similar approach for investigating the knowle…

Studying and improving reasoning in humans and machines

2023-09-21 · Nicolas Yax, Hernan Anlló, Stefano Palminteri

In the present study, we investigate and compare reasoning in large language models (LLM) and humans using a selection of cognitive psychology tools traditionally dedicated to the study of (bounded) rationality. To do so…

Cognitive phantoms in LLMs through the lens of latent variables

2024-09-06 · Sanne Peereboom, Inga Schwabe, Bennett Kleinberg

Large language models (LLMs) increasingly reach real-world applications, necessitating a better understanding of their behaviour. Their size and complexity complicate traditional assessment methods, causing the emergence…

On the Cognition of Visual Question Answering Models and Human Intelligence: A Comparative Study

2023-10-04 · Liben Chen, Long Chen, Tian Ellison-Chen, Zhuoyuan Xu

Visual Question Answering (VQA) is a challenging task that requires cross-modal understanding and reasoning of visual image and natural language question. To inspect the association of VQA models to human cognition, we d…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)