paper-with-me

Papers

Exploring Vision Language Models for Facial Attribute Recognition: Emotion, Race, Gender, and Age

2024-10-31 · Nouar AlDahoul, Myles Joshua Toledo Tan, Harishwar Reddy Kasireddy, Yasir Zaki

Technologies for recognizing facial attributes like race, gender, age, and emotion have several applications, such as surveillance, advertising content, sentiment analysis, and the study of demographic trends and social behaviors. Analyzing demographic characteristics based on images and analyzing facial expressions have several challenges due to the complexity of humans' facial attributes. Traditional approaches have employed CNNs and various other deep learning techniques, trained on extensive collections of labeled images. While these methods demonstrated effective performance, there remains potential for further enhancements. In this paper, we propose to utilize vision language models (VLMs) such as generative pre-trained transformer (GPT), GEMINI, large language and vision assistant (LLAVA), PaliGemma, and Microsoft Florence2 to recognize facial attributes such as race, gender, age, and emotion from images with human faces. Various datasets like FairFace, AffectNet, and UTKFace have been utilized to evaluate the solutions. The results show that VLMs are competitive if not superior to traditional techniques. Additionally, we propose "FaceScanPaliGemma"--a fine-tuned PaliGemma model--for race, gender, age, and emotion recognition. The results show an accuracy of 81.1%, 95.8%, 80%, and 59.4% for race, gender, age group, and emotion classification, respectively, outperforming pre-trained version of PaliGemma, other VLMs, and SotA methods. Finally, we propose "FaceScanGPT", which is a GPT-4o model to recognize the above attributes when several individuals are present in the image using a prompt engineered for a person with specific facial and/or physical attributes. The results underscore the superior multitasking capability of FaceScanGPT to detect the individual's attributes like hair cut, clothing color, postures, etc., using only a prompt to drive the detection and recognition tasks.

📄 PDF Abstract BibTeX arXiv:2410.24148

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeEmotion ClassificationEmotion RecognitionSentiment Analysis

Similar Papers 제목 키워드 기반

Exp-Graph: How Connections Learn Facial Attributes in Graph-based Expression Recognition

2025-07-19 · Nandani Sharma, Dinesh Singh arxiv

Facial expression recognition is crucial for human-computer interaction applications such as face animation, video surveillance, affective computing, medical analysis, etc. Since the structure of facial attributes varies…

Facial Expression Recognition

The Good, the Better, and the Best: Improving the Discriminability of Face Embeddings through Attribute-aware Learning

2026-03-16 · Ana Dias, João Ribeiro Pinto, Hugo Proença, João C. Neves arxiv

Despite recent advances in face recognition, robust performance remains challenging under large variations in age, pose, and occlusion. A common strategy to address these issues is to guide representation learning with a…

Representation LearningFace VerificationFace Recognition

Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning

2025-04-09 · Ashutosh Chaubey, Xulang Guan, Mohammad Soleymani

The human face plays a central role in social communication, necessitating the use of performant computer vision tools for human-centered applications. We propose Face-LLaVA, a multimodal large language model for face-ce…

Action Unit DetectionAge EstimationAttributeDeepFake Detection+5

Exploring Correlations in Multiple Facial Attributes through Graph Attention Network

2018-10-22 · Yan Zhang, Li Sun

Estimating multiple attributes from a single facial image gives comprehensive descriptions on the high level semantics of the face. It is naturally regarded as a multi-task supervised learning problem with a single deep …

AttributeGraph AttentionMulti-Task Learning

Leveraging vision-language models for fair facial attribute classification

2024-03-15 · Miao Zhang, Rumi Chunara

Performance disparities of image recognition across different demographic populations are known to exist in deep learning-based models, but previous work has largely addressed such fairness problems assuming knowledge of…

AttributeFacial Attribute ClassificationFairnessLanguage Modeling+1