paper-with-me

Papers

Dataset Scale and Societal Consistency Mediate Facial Impression Bias in Vision-Language AI

2024-08-04 · Robert Wolfe, Aayushi Dangol, Alexis Hiniker, Bill Howe

Multimodal AI models capable of associating images and text hold promise for numerous domains, ranging from automated image captioning to accessibility applications for blind and low-vision users. However, uncertainty about bias has in some cases limited their adoption and availability. In the present work, we study 43 CLIP vision-language models to determine whether they learn human-like facial impression biases, and we find evidence that such biases are reflected across three distinct CLIP model families. We show for the first time that the the degree to which a bias is shared across a society predicts the degree to which it is reflected in a CLIP model. Human-like impressions of visually unobservable attributes, like trustworthiness and sexuality, emerge only in models trained on the largest dataset, indicating that a better fit to uncurated cultural data results in the reproduction of increasingly subtle social biases. Moreover, we use a hierarchical clustering approach to show that dataset size predicts the extent to which the underlying structure of facial impression bias resembles that of facial impression bias in humans. Finally, we show that Stable Diffusion models employing CLIP as a text encoder learn facial impression biases, and that these biases intersect with racial biases in Stable Diffusion XL-Turbo. While pretrained CLIP models may prove useful for scientific studies of bias, they will also require significant dataset curation when intended for use as general-purpose models in a zero-shot setting.

📄 PDF Abstract BibTeX arXiv:2408.01959

Code (0)

등록된 구현이 없습니다.

Tasks

Image Captioning

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

StoryMaker: Towards Holistic Consistent Characters in Text-to-image Generation

2024-09-19 · Zhengguang Zhou, Jing Li, Huaxia Li, Nemo Chen 외

Tuning-free personalized image generation methods have achieved significant success in maintaining facial consistency, i.e., identities, even with multiple characters. However, the lack of holistic consistency in scenes …

Image GenerationPersonalized Image GenerationText to Image GenerationText-to-Image Generation

SelFSR: Self-Conditioned Face Super-Resolution in the Wild via Flow Field Degradation Network

2021-12-20 · Xianfang Zeng, Jiangning Zhang, Liang Liu, Guangzhong Tian 외

In spite of the success on benchmark datasets, most advanced face super-resolution models perform poorly in real scenarios since the remarkable domain gap between the real images and the synthesized training pairs. To ta…

Super-Resolution

Face-MakeUpV2: Facial Consistency Learning for Controllable Text-to-Image Generation

2025-10-17 · Dawei Dai, Yinxiu Zhou, Chenghang Li, Guolai Jiang 외 arxiv

In facial image generation, current text-to-image models often suffer from facial attribute leakage and insufficient physical consistency when responding to local semantic instructions. In this study, we propose Face-Mak…

Text-to-Image Generation

Unsupervised Domain Attention Adaptation Network for Caricature Attribute Recognition

2020-07-18 · ECCV 2020 8 · Wen Ji, Kelei He, Jing Huo, Zheng Gu 외

Caricature attributes provide distinctive facial features to help research in Psychology and Neuroscience. However, unlike the facial photo attribute datasets that have a quantity of annotated images, the annotations of …

AttributeCaricatureDomain AdaptationUnsupervised Domain Adaptation

FFHQ-Makeup: Paired Synthetic Makeup Dataset with Facial Consistency Across Multiple Styles

2025-08-05 · Xingchao Yang, Shiori Ueda, Yuantian Huang, Tomoya Akiyama 외 arxiv

Paired bare-makeup facial images are essential for a wide range of beauty-related tasks, such as virtual try-on, facial privacy protection, and facial aesthetics analysis. However, collecting high-quality paired makeup d…

Text-to-Image GenerationVirtual Try-on