paper-with-me

Papers

Beyond Vision: How Large Language Models Interpret Facial Expressions from Valence-Arousal Values

2025-02-08 · Vaibhav Mehra, Guy Laban, Hatice Gunes

Large Language Models primarily operate through text-based inputs and outputs, yet human emotion is communicated through both verbal and non-verbal cues, including facial expressions. While Vision-Language Models analyze facial expressions from images, they are resource-intensive and may depend more on linguistic priors than visual understanding. To address this, this study investigates whether LLMs can infer affective meaning from dimensions of facial expressions-Valence and Arousal values, structured numerical representations, rather than using raw visual input. VA values were extracted using Facechannel from images of facial expressions and provided to LLMs in two tasks: (1) categorizing facial expressions into basic (on the IIMI dataset) and complex emotions (on the Emotic dataset) and (2) generating semantic descriptions of facial expressions (on the Emotic dataset). Results from the categorization task indicate that LLMs struggle to classify VA values into discrete emotion categories, particularly for emotions beyond basic polarities (e.g., happiness, sadness). However, in the semantic description task, LLMs produced textual descriptions that align closely with human-generated interpretations, demonstrating a stronger capacity for free text affective inference of facial expressions.

📄 PDF Abstract BibTeX arXiv:2502.06875

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

SemanticFace: Semantic Facial Action Estimation via Semantic Distillation in Interpretable Space

2026-03-16 · Zejian Kang, Kai Zheng, Yuanchen Fei, Wentao Yang 외 arxiv

Facial action estimation from a single image is often formulated as predicting or fitting parameters in compact expression spaces, which lack explicit semantic interpretability. However, many practical applications, such…

Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond

2025-11-01 · Fan Zhang, Haoxuan Li, Shengju Qian, Xin Wang 외 arxiv

Multimodal Large Language Models (MLLMs) have revolutionized numerous research fields, including computer vision and affective computing. As a pivotal challenge in this interdisciplinary domain, facial expression recogni…

Facial Expression RecognitionReinforcement Learning

CLIPER: A Unified Vision-Language Framework for In-the-Wild Facial Expression Recognition

2023-03-01 · Hanting Li, Hongjing Niu, Zhaoqing Zhu, Feng Zhao

Facial expression recognition (FER) is an essential task for understanding human behaviors. As one of the most informative behaviors of humans, facial expressions are often compound and variable, which is manifested by t…

Dynamic Facial Expression RecognitionFacial Expression RecognitionFacial Expression Recognition (FER)

KeyframeFace: Language-Driven Facial Animation via Semantic Keyframes

2025-12-12 · Jingchao Wu, Zejian Kang, Haibo Liu, Yuanchen Fei 외 arxiv

Facial animation is a core component for creating digital characters in Computer Graphics (CG) industry. A typical production workflow relies on sparse, semantically meaningful keyframes to precisely control facial expre…

A Classification Model Utilizing Facial Landmark Tracking to Determine Sentence Types for American Sign Language Recognition

2022-11-23 · Janice Nguyen, Y. Curtis Wang

The deaf and hard of hearing community relies on American Sign Language (ASL) as their primary mode of communication, but communication with others who do not know ASL can be difficult, especially during emergencies wher…

Landmark TrackingSentenceSign Language Recognition