paper-with-me

홈 › Papers

Social perception of faces in a vision-language model

2024-08-26 · Carina I. Hausladen, Manuel Knott, Colin F. Camerer, Pietro Perona

We explore social perception of human faces in CLIP, a widely used open-source vision-language model. To this end, we compare the similarity in CLIP embeddings between different textual prompts and a set of face images. Our textual prompts are constructed from well-validated social psychology terms denoting social perception. The face images are synthetic and are systematically and independently varied along six dimensions: the legally protected attributes of age, gender, and race, as well as facial expression, lighting, and pose. Independently and systematically manipulating face attributes allows us to study the effect of each on social perception and avoids confounds that can occur in wild-collected data due to uncontrolled systematic correlations between attributes. Thus, our findings are experimental rather than observational. Our main findings are three. First, while CLIP is trained on the widest variety of images and texts, it is able to make fine-grained human-like social judgments on face images. Second, age, gender, and race do systematically impact CLIP's social perception of faces, suggesting an undesirable bias in CLIP vis-a-vis legally protected attributes. Most strikingly, we find a strong pattern of bias concerning the faces of Black women, where CLIP produces extreme values of social perception across different ages and facial expressions. Third, facial expression impacts social perception more than age and lighting as much as age. The last finding predicts that studies that do not control for unprotected visual attributes may reach the wrong conclusions on bias. Our novel method of investigation, which is founded on the social psychology literature and on the experiments involving the manipulation of individual attributes, yields sharper and more reliable observations than previous observational methods and may be applied to study biases in any vision-language model.

📄 PDF Abstract BibTeX arXiv:2408.14435

Code (1)

carinahausladen/clip-face-bias 공식 구현

Tasks

Language ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Neural encoding of real world face perception

2025-05-13 · Arish Alreja, Michael J. Ward, Lisa S. Parker, R. Mark Richardson 외

Social perception unfolds as we freely interact with people around us. We investigated the neural basis of real world face perception using multi electrode intracranial recordings in humans during spontaneous interaction…

Social-LLaVA: Enhancing Robot Navigation through Human-Language Reasoning in Social Spaces

2024-12-30 · Amirreza Payandeh, Daeun Song, Mohammad Nazeri, Jing Liang 외

Most existing social robot navigation techniques either leverage hand-crafted rules or human demonstrations to connect robot perception to socially compliant actions. However, there remains a significant gap in effective…

2kRobot NavigationVisual Question Answering (VQA)

Learning to see people like people

2017-05-05 · Amanda Song, Linjie Li, Chad Atalla, Garrison Cottrell

Humans make complex inferences on faces, ranging from objective properties (gender, ethnicity, expression, age, identity, etc) to subjective judgments (facial attractiveness, trustworthiness, sociability, friendliness, e…

Behavioral Geometric Supervision Aligns Video Foundation Models with Human Social Perception

2025-10-01 · Kathy Garcia, Leyla Isik arxiv

Current video foundation models, including the strongest self-supervised models such as V-JEPA2, fail to capture how humans organize social information in dynamic scenes. For example, across a range of diverse vision mod…

FaceShield: Explainable Face Anti-Spoofing with Multimodal Large Language Models

2025-05-14 · Hongyang Wang, Yichen Shi, Zhuofu Tao, Yuhao Gao 외

Face anti-spoofing (FAS) is crucial for protecting facial recognition systems from presentation attacks. Previous methods approached this task as a classification problem, lacking interpretability and reasoning behind th…

Face Anti-Spoofing