paper-with-me

Papers

MLLM-based Textual Explanations for Face Comparison

2026-03-17 · Redwan Sony, Anil K Jain, Arun Ross arxiv

Multimodal Large Language Models (MLLMs) have recently been proposed as a means to generate natural-language explanations for face recognition decisions. While such explanations facilitate human interpretability, their reliability on unconstrained face images remains underexplored. In this work, we systematically analyze MLLM-generated explanations for the unconstrained face verification task on the challenging IJB-S dataset, with a particular focus on extreme pose variation and surveillance imagery. Our results show that even when MLLMs produce correct verification decisions, the accompanying explanations frequently rely on non-verifiable or hallucinated facial attributes that are not supported by visual evidence. We further study the effect of incorporating information from traditional face recognition systems, viz., scores and decisions, alongside the input images. Although such information improves categorical verification performance, it does not consistently lead to faithful explanations. To evaluate the explanations beyond decision accuracy, we introduce a likelihood-ratio-based framework that measures the evidential strength of textual explanations. Our findings highlight fundamental limitations of current MLLMs for explainable face recognition and underscore the need for a principled evaluation of reliable and trustworthy explanations in biometric applications. Code is available at https://github.com/redwankarimsony/LR-MLLMFR-Explainability.

📄 PDF Abstract BibTeX arXiv:2603.16629

Code (0)

등록된 구현이 없습니다.

Tasks

Face VerificationFace Recognition

Similar Papers 제목 키워드 기반

Fake-in-Facext: Towards Fine-Grained Explainable DeepFake Analysis

2025-10-23 · Lixiong Qin, Yang Zhang, Mei Wang, Jiani Hu 외 arxiv

The advancement of Multimodal Large Language Models (MLLMs) has bridged the gap between vision and language tasks, enabling the implementation of Explainable DeepFake Analysis (XDFA). However, current methods suffer from…

Multi-Task Learning

Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

2026-08-07 · Xuechao Zou, Shun Zhang, Kai Li, Yi Zhou 외 hf

The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety. However, existing deepfake video benchmark…

Face Swapping

Unlocking the Forgery Detection Potential of Vanilla MLLMs: A Novel Training-Free Pipeline

2025-11-17 · Rui Zuo, Qinyue Tong, Zhe-Ming Lu, Ziqian Lu arxiv

With the rapid advancement of artificial intelligence-generated content (AIGC) technologies, including multimodal large language models (MLLMs) and diffusion models, image generation and manipulation have become remarkab…

Image Generation

AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs

2026-01-06 · Boyu Chang, Qi Wang, Xi Guo, Zhixiong Nan 외 arxiv

Visual abductive reasoning (VAR) is a challenging task that requires AI systems to infer the most likely explanation for incomplete visual observations. While recent MLLMs develop strong general-purpose multimodal reason…

Multimodal Reasoning

Guard Me If You Know Me: Protecting Specific Face-Identity from Deepfakes

2025-05-26 · Kaiqing Lin, Zhiyuan Yan, Ke-Yue Zhang, Li Hao 외

Securing personal identity against deepfake attacks is increasingly critical in the digital age, especially for celebrities and political figures whose faces are easily accessible and frequently targeted. Most existing d…

DeepFake DetectionFace GenerationFace SwappingLarge Language Model+1