paper-with-me

Papers

FaceInsight: A Multimodal Large Language Model for Face Perception

2025-04-22 · Jingzhi Li, Changjiang Luo, Ruoyu Chen, Hua Zhang, Wenqi Ren, Jianhou Gan, Xiaochun Cao

Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perform poorly in face perception tasks, often producing inaccurate or misleading responses to face-specific queries. To address this gap, we propose FaceInsight, the versatile face perception MLLM that provides fine-grained facial information. Our approach introduces visual-textual alignment of facial knowledge to model both uncertain dependencies and deterministic relationships among facial information, mitigating the limitations of language-driven reasoning. Additionally, we incorporate face segmentation maps as an auxiliary perceptual modality, enriching the visual input with localized structural cues to enhance semantic understanding. Comprehensive experiments and analyses across three face perception tasks demonstrate that FaceInsight consistently outperforms nine compared MLLMs under both training-free and fine-tuned settings.

📄 PDF Abstract BibTeX arXiv:2504.15624

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model

Similar Papers 제목 키워드 기반

Face-MLLM: A Large Face Perception Model

2024-10-28 · Haomiao Sun, Mingjie He, Tianheng Lian, Hu Han 외

Although multimodal large language models (MLLMs) have achieved promising results on a wide range of vision-language tasks, their ability to perceive and understand human faces is rarely explored. In this work, we compre…

AttributemodelQuestion AnsweringVisual Question Answering

A Modern System Recipe for Situated Embodied Human-Robot Conversation with Real-Time Multimodal LLMs and Tool-Calling

2026-02-04 · Dong Won Lee, Sarah Gillet, Louis-Philippe Morency, Cynthia Breazeal 외 arxiv

Situated embodied conversation requires robots to interleave real-time dialogue with active perception: deciding what to look at, when to look, and what to say under tight latency constraints. We present a simple, minima…

Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention

2026-05-21 · Changyuan Tian, Zhicong Lu, Huaxing Liu, Xiang Wang 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising paradigm for advancing complex reasoning in large language models, and recent work extends RLVR to multimodal large language models (MLLMs)…

Reinforcement LearningMultimodal Reasoning

FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMs

2025-03-27 · CVPR 2025 1 · Xiaoqin Wang, Xusen Ma, Xianxu Hou, Meidan Ding 외

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in various tasks. However, effectively evaluating these MLLMs on face perception remains largely unexplored. To address this gap, we intr…

AttributeBenchmarkingQuestion AnsweringVisual Question Answering+1

AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception

2024-01-16 · Yipo Huang, Quan Yuan, Xiangfei Sheng, Zhichao Yang 외

With collective endeavors, multimodal large language models (MLLMs) are undergoing a flourishing development. However, their performances on image aesthetics perception remain indeterminate, which is highly desired in re…

MLLM Evaluation: Aesthetics